Pith. sign in

Paper Citation Record · LEDGER

CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2106.11097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.11097 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:57:09.085648Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T04:48:54.299822Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d75c229d-e26a-4636-9a92-0421d108cc6f · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.683764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:31c2b5b92e56f5ef70f179881c4d0321f722018f8a1629fa82cdf5b78e3d4b58

Observation 4a47bb5f-7b6b-42c4-8b69-3b555bf773fd · inbound

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels cites this paper.

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:48:54.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:46:18.542521Z digest=sha256:fed3ebbf53813063fc0e4b51dc8f33ac06ffbeafeb8d4206aec0cae43f8bb6b6

Observation e210b03a-5f04-4acf-8706-2e12d001248a · inbound

Detecting Content Rating Violations in Android Applications: A Vision-Language Approach cites this paper.

Detecting Content Rating Violations in Android Applications: A Vision-Language Approach CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T21:57:09.085648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:57:09.085648Z digest=sha256:44399e9cb1dc3781be0864ea8da86a1424997b841133a3f28b75a9cb915f6a69

Observation b372a1b7-fd12-4298-83af-85acfe7bdb65 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.297596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:ef9d570f137cb0e4dc4d1e718682b5135aec7744c0e87c45d9db1f4bc41c274b

Observation 06498fb3-0c1d-4cf3-921d-2ca0b4ee5cfb · inbound

A Mathematical Perspective On Contrastive Learning cites this paper.

A Mathematical Perspective On Contrastive Learning CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.738221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.738221Z digest=sha256:d0041faece3d479bc7937621597da46718dba676fe7e99d502259cf6401d2b63

Observation ddea426d-f18b-44c0-99df-70fe5494a31e · inbound

Learning Speaker-Invariant Visual Features for Lipreading cites this paper.

Learning Speaker-Invariant Visual Features for Lipreading CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:57.052410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:57.052410Z digest=sha256:87c175171ccb0cac4777bcbc57a4f96dcbc7cb3b82ea3a47e3f2d3e53540a03d

Observation 475e19fe-8987-4049-afcf-545444f1c113 · inbound

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval cites this paper.

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:36:27.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:36:27.471308Z digest=sha256:b61dc7e8cb376dee41fd2bc0b58d13bbf9cf14864b40463c0c121c0b5de59fba

Observation 8adc1884-5841-4d92-9623-872b8dd0785b · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.766732Z digest=sha256:675eaf2eca37aabed5ba5149f1e30d014b287f71f43be20e5291153466b70820

Observation 18390aa0-4e10-4ff7-94b5-3a14c7a0bbbc · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.914340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:61e7516fdd0a4d4ad4c979e02ee7c562c1a41ffcce67d51781fdba5103de0a64

Observation 7bc0539d-8372-4464-8dcb-3e49fe2c6d7f · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.665420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:3c7452568888ae7fb5abccc120e2c706ca3d4ddbc2762557e61d79ab9803fa72

Observation 15cdf844-dbd2-4316-8418-c8d476152921 · inbound

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation cites this paper.

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.314300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T20:04:39.623342Z digest=sha256:313c306abfc5b5f4f509e5aa11d3529a58854d4733b2ac9f7e4c3f5c41b30484

Observation e7f17c4f-79fc-4248-98ae-e47fbf162147 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.640974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:7b87eaf39cd27b97be3977c8c645b9eade92b06b7f4cc18710c61e9eba987519

Observation 599e28ca-8d4e-46ce-9590-96fdadc8c36b · inbound

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval cites this paper.

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T04:19:16.924621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T04:19:16.924621Z digest=sha256:598d0b1c81c1b3f252c678ff55d134f8931f612553bb077fb3cc490ae7b3307c

Observation 8de72bdb-d372-4f2c-9b39-7a4f0850d79a · inbound

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification cites this paper.

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:00.942499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:01:00.942499Z digest=sha256:c388437ce18f1d1815b575314ccfb86e5f0c4b9a20eff0bdf7e3abf8549ce84d