Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2209.06430.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:58.251948Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T15:17:07.354594Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8658039f-ff8d-44e3-b8bb-05a068784109 · inbound
A Survey on Foundation Models for Personalized Federated Intelligence CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d526405-4e77-4079-a62e-68b17786e9ee · inbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d61d9329-9b4b-4dd8-8ca8-50718ab3065b · inbound
Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68df01c5-db36-4869-a1ba-cc133ae62b8f · inbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ca38b7-a66a-469a-9522-af7642c6a203 · inbound
Adversarial Video Promotion Against Text-to-Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 31c97ae3-2ec7-4ac8-b0bc-0bd617d8d8d7 · inbound
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d93407-97c0-4006-adbd-0514b02ba657 · inbound
Adapting MLLMs for Nuanced Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de5b80b0-bbc7-4067-a663-182999067c1c · inbound
CoVR-R:Reason-Aware Composed Video Retrieval CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e6ccf5-b049-414e-b36a-35f80997fd69 · inbound
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c807420-93b1-4ad1-948b-697b97b30293 · inbound
Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0bf0880-0c9a-4610-a8d4-1c18e2e63f2b · inbound
Knowledge-guided Disentanglement with Atomic Actions for Action Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.