Pith. sign in

Paper Citation Record · LEDGER

CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2209.06430.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2209.06430 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:58.251948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:17:07.354594Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8658039f-ff8d-44e3-b8bb-05a068784109 · inbound

A Survey on Foundation Models for Personalized Federated Intelligence cites this paper.

A Survey on Foundation Models for Personalized Federated Intelligence CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:34:57.747854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T15:32:15.293888Z digest=sha256:20c45240cc7d9674b67966b63c3fdccd0fa6f8c0eea6704edbb4167e58f259c1

Observation 7d526405-4e77-4079-a62e-68b17786e9ee · inbound

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition cites this paper.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.251948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.251948Z digest=sha256:4e455c3c56886ab0efa6f30007925178d35bcb64a7f589ec43f026a58fcd4ee6

Observation d61d9329-9b4b-4dd8-8ca8-50718ab3065b · inbound

Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos cites this paper.

Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:07.541908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:07.541908Z digest=sha256:0532a3ecbe2e98b9f05c899f1a0e77a372365177efa55a28541ca6c824bee34a

Observation 68df01c5-db36-4869-a1ba-cc133ae62b8f · inbound

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning cites this paper.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.684895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.684895Z digest=sha256:6645822b59cf35559de022b71a644fd68d37f0efb056956ff48ccfd6bda97cb5

Observation 20ca38b7-a66a-469a-9522-af7642c6a203 · inbound

Adversarial Video Promotion Against Text-to-Video Retrieval cites this paper.

Adversarial Video Promotion Against Text-to-Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.131999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:05:07.182361Z digest=sha256:74bca8b989c8b9c0d97d07a8a34e05faa6ca9f039059456371a9dddcf4ea271d

Observation 31c97ae3-2ec7-4ac8-b0bc-0bd617d8d8d7 · inbound

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications cites this paper.

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:58.254722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:58.254722Z digest=sha256:253f3903c0c6746d8f3b4059993022a8422d0216a53a8cc67551be3414705722

Observation 56d93407-97c0-4006-adbd-0514b02ba657 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.941379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:ddd2bab744ec17a2a0b9d0a85003824f157502a52a6b6eba5904c0ae473efe6e

Observation de5b80b0-bbc7-4067-a663-182999067c1c · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:6aec53475fc0c3ce08b91ea55adc08d47a64d77df15f877e290b0ea363d0eabd

Observation 22e6ccf5-b049-414e-b36a-35f80997fd69 · inbound

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement cites this paper.

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:17:07.356298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T15:09:47.855795Z digest=sha256:af9df80d44c430fca083f133c78378d6d224a816402a499ba7bcdcf82b990bc6

Observation 7c807420-93b1-4ad1-948b-697b97b30293 · inbound

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing cites this paper.

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T08:55:44.783824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T08:55:44.783824Z digest=sha256:0d04400dad519f62862a7df27acb769251df6b1ce8d76d619f7a24aec030411e

Observation e0bf0880-0c9a-4610-a8d4-1c18e2e63f2b · inbound

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition cites this paper.

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:42.840512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:02:42.840512Z digest=sha256:495f81bdad925bb129a0fd23ef2fc19de6c76db9e4c2338fd36b35fc77af0b93