Pith. sign in

Paper Citation Record · LEDGER

CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2106.11097.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.11097 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:57:09.085648Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T04:48:54.299822Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d75c229d-e26a-4636-9a92-0421d108cc6f · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.683764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:825e7bfe0a53d4507dc526667dc953b0e8e49ba62cd1699d77a1cbf71c7130e2

Observation 4a47bb5f-7b6b-42c4-8b69-3b555bf773fd · inbound

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels cites this paper.

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:48:54.303371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:46:18.542521Z digest=sha256:e357822d7f7ea4b0b497328b4e7cd149b4d0bb0b7048031104bf6a3b6099ba71

Observation e210b03a-5f04-4acf-8706-2e12d001248a · inbound

Detecting Content Rating Violations in Android Applications: A Vision-Language Approach cites this paper.

Detecting Content Rating Violations in Android Applications: A Vision-Language Approach CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T21:57:09.085648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:57:09.085648Z digest=sha256:feff16aca519829b90d2afbbfd1affc0115a126d126e7f8a1b423f0523f2161d

Observation b372a1b7-fd12-4298-83af-85acfe7bdb65 · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.297596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:0d3856192a1cd66e205a1d2d1252249ab7e4b673fff6383ec3fc9ea88fb3edf9

Observation 06498fb3-0c1d-4cf3-921d-2ca0b4ee5cfb · inbound

A Mathematical Perspective On Contrastive Learning cites this paper.

A Mathematical Perspective On Contrastive Learning CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:09.738221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:09.738221Z digest=sha256:147136550a3bfdee71dab9885d44489641d13b86b37f888eb0b0180f62aeb58a

Observation ddea426d-f18b-44c0-99df-70fe5494a31e · inbound

Learning Speaker-Invariant Visual Features for Lipreading cites this paper.

Learning Speaker-Invariant Visual Features for Lipreading CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:57.052410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:57.052410Z digest=sha256:e72b2d2ffd2defe2ce4e04d851cf6c22f27b7566d80cd0b29be68e7667a819ab

Observation 475e19fe-8987-4049-afcf-545444f1c113 · inbound

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval cites this paper.

Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:36:27.471308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:36:27.471308Z digest=sha256:89896d7c9fcef86838aef05d0147dc1e06d8a12601062b83b1f2b2a01c6a64e0

Observation 8adc1884-5841-4d92-9623-872b8dd0785b · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:39.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:39.766732Z digest=sha256:7089788c041b77579a9fc23e3d33f1fa79cc2924ea6a7f4709fea10d5a7ce85a

Observation 18390aa0-4e10-4ff7-94b5-3a14c7a0bbbc · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.914340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:53:51.162967Z digest=sha256:e0cbf5e4f99ca6c1e0b05b1b3adb4da730cbc81f8fd47fc1dfea2b6930356a0b

Observation 7bc0539d-8372-4464-8dcb-3e49fe2c6d7f · inbound

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models cites this paper.

EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.665420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:16:15.202466Z digest=sha256:54e01cdf635d0a1257b1ccf30719b9583667c67ade79fe5ee28fc1c16aa0f5f5

Observation 15cdf844-dbd2-4316-8418-c8d476152921 · inbound

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation cites this paper.

Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.314300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T20:04:39.623342Z digest=sha256:28d7f2f73cb38a9fa8489018b27a893bb8383fc5df316069344aa7430870bb43

Observation e7f17c4f-79fc-4248-98ae-e47fbf162147 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.640974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:89c227be962f76fdab3a589772ff8a2a0e34275905ddf92f0c1b4052ecec5d0f

Observation 599e28ca-8d4e-46ce-9590-96fdadc8c36b · inbound

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval cites this paper.

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T04:19:16.924621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T04:19:16.924621Z digest=sha256:159e548d8858be3ec09eb62e53e3b2fa929393ad61e9ca49939649ebab638c36

Observation 8de72bdb-d372-4f2c-9b39-7a4f0850d79a · inbound

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification cites this paper.

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:00.942499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:01:00.942499Z digest=sha256:f3809d273d8aebba79d8f6e696ba234e87f3d6a108744399267cac0e0d258627