Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2107.06383.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T00:56:16.603472Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T00:36:53.284895Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 06c965ac-515d-47b0-8d86-5a13fd5e9f36 · inbound
Hierarchical Text-Conditional Image Generation with CLIP Latents How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 63e1a544-2b76-4f2d-8f94-10bb981776a0 · inbound
CoCa: Contrastive Captioners are Image-Text Foundation Models How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 257b8609-2703-419d-bd7f-42cc87b29625 · inbound
GIT: A Generative Image-to-text Transformer for Vision and Language How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13ebeee1-ad8b-49bf-9ba1-1854afecb28b · inbound
LAION-5B: An open large-scale dataset for training next generation image-text models How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb05c3d7-b7f6-4b6a-aace-20d7ebb73939 · inbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7825d19f-eb53-4f0d-bfd3-6f58d4a5bcf9 · inbound
VideoChat: Chat-Centric Video Understanding How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c46e81a1-d775-4997-a9d0-c1b7f6d1e4f4 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c908abb-08bd-4fef-9f0e-19a196ab2a6e · inbound
GPT-4V(ision) is a Generalist Web Agent, if Grounded How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f51676e-46bc-46ab-a39e-c9b364ed7332 · inbound
Visual Question Answering on Multiple Remote Sensing Image Modalities How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a20d7a-f05e-4890-8a86-9b15808cf933 · inbound
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 755b5c06-7c0f-48b4-8ccb-8c29caa83c8e · inbound
(Almost) Free Modality Stitching of Foundation Models How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed88d33-cd88-4c23-a33b-68593143e07b · inbound
LSDM: LLM-Enhanced Spatio-temporal Diffusion Model for Service-Level Mobile Traffic Prediction How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb8275f-7210-4586-be21-74a5ec661e98 · inbound
Multi-Agent Cooperative Learning for Robust Vision-Language Alignment under OOD Concepts How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a78330c6-8c21-4bbd-967d-f7cc685ce06c · inbound
A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fb61f7-08b2-490e-be4c-f1f62cb27ab2 · inbound
ReVision : A Post-Hoc, Vision-Based Technique for Replacing Unacceptable Concepts in Image Generation Pipeline How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5663d9a4-57a4-438f-9d8d-96596ad24329 · inbound
Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5b00289-6703-4814-b945-2178383806fe · inbound
Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3be7cf25-66f4-452c-ad79-910c97532f0b · inbound
UniMesh: Unifying 3D Mesh Understanding and Generation How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8bb81e6-b1e9-46d1-9b9b-5215e824a818 · inbound
MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 808ed0c4-435b-46a9-a0b0-dbec6a5eae76 · inbound
Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis How Much Can CLIP Benefit Vision-and-Language Tasks?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.