Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:57:12.054681Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 9 inbound Pith citation observations for arXiv:2411.16828.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:57:12.054681Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:29:57.356586Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:39:50.707139Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 87a9bce5-8af8-4d75-aece-df6652e9943c · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d27400c-8ca4-4ae7-9221-91b6b7517cc0 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions nocaps: novel object caption- ing at scale
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9c7aa08c-03a6-4128-be0d-7933828bdbe0 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a64e86-c10a-46ae-b785-ac4b59bec140 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Contrastive Localized Language-Image Pre-Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d865697b-4c7d-4f3d-8ded-4c99fe4736f3 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa93cab-c879-401a-93bf-f6af44de8b9a · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c63e58-115e-4a5c-9fcf-14d6cd4b683c · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d3cbf0-4c91-43ab-b5ec-87d41022b000 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Uniter: Universal image-text representation learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 856507f4-b276-45cd-b87d-591fca256d3d · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19eac34b-3930-450a-a17e-cc60b3b72e55 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8190cd5-b06b-4614-854d-11611a82aa61 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Imagenet: A large-scale hierarchical image database
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ebdab8-4d15-4188-a959-e3da8492db67 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f5d061-4bfd-406c-acde-9f0c709a33cf · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a20a60-25b5-4c7f-a116-bb3244b52271 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improving clip training with language rewrites
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 23714bab-d68e-4368-b4ad-f8eb6dbd2213 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Data Filtering Networks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d164bd10-f6a1-4f7e-a598-75097867be19 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 675b36ef-7343-423b-a2b9-774c7b581dac · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Dat- acomp: In search of the next generation of multimodal datasets
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c0e496eb-6e35-46fc-8f0d-99aa2501f4af · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e9be35-437e-48ee-be88-39a63aab6713 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7e1d1dc0-b2cb-4289-961c-7305b5a4e58e · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Open- clip
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4e663a27-f256-4864-8de8-a8110ad34162 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77d857e-3b87-4a6d-9e35-25bcd6e60d66 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Vilt: Vision- and-language transformer without convolution or region su- pervision
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fafa437-3385-432e-ad67-33adf7348b28 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Veclip: Improving clip training via visual-enriched captions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9299736a-f2b8-4ea5-9d26-fa4dbc5a0f44 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Modeling Caption Diversity in Contrastive Vision-Language Pretraining
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c672f7-5eae-4a63-a2d5-3b53ef73108d · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Align before fuse: Vision and language representation learn- ing with momentum distillation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46fb3ea-2ed2-4ca7-a8c7-00abf378e968 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fed288c3-f5a6-467a-aa17-970573b7dd0f · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1171e4b-2022-4e26-9bcf-245de887639f · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4886e482-f71b-40e5-baa4-cff250d1c4af · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b79ca7-7d68-4d30-ae59-f2fc638f8802 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions What If We Recaption Billions of Web Images with LLaMA-3?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118f6ee0-5eb3-442b-834f-a3b6a43de51f · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions An inverse scal- ing law for clip training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3d6e3517-fe59-4b9f-8675-3d1cd0181484 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Evaluating object hallucination in large vision-language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e13e9a0-7ba8-4341-8a23-3029afa06591 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b20fcd7b-2378-4c47-be2c-be2e2a52e369 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Microsoft coco: Common objects in context
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c7bcbe9-72e3-4cc9-a67d-2f9137696731 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improved baselines with visual instruction tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 974a6fed-75b7-42fe-a859-4a80f5947e52 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Visual instruction tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2def70fc-4104-4b95-9f14-89fd72aaa89d · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MLLMs-Augmented Visual-Language Representation Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79aad07b-cc57-4341-828f-90a2f2aff7a7 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b816f311-eb6c-47b7-8a25-e3285280ef85 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1f25dc0d-99d3-4021-8a8e-2ab38f1d389f · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Introducing meta llama 3: The most capable openly available llm to date
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 95029b9b-1135-47f5-9d57-a8d123894450 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improving multimodal datasets with image captioning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 864f0c83-36d8-4123-8902-435d856882bb · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Representation Learning with Contrastive Predictive Coding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d418dd9-57d1-4748-8765-a003e296d0f6 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Introducing chatgpt
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7bc454e2-6576-4262-a68c-d1ddbcdcf4b4 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf99a94-d00c-45e2-846b-2441208eb8da · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Learning transferable visual models from natural language supervi- sion
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d2d1b4a7-cc30-46b9-b26b-46a15f778456 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 717c274e-ca7c-44aa-9a32-90881f0de1df · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Towards vqa models that can read
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35637afb-a9e7-457d-99da-76ec8784bc20 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d8bc428a-7b52-45df-a683-b0a6d8981a14 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MOFI: Learning Image Representations from Noisy Entity Annotated Images
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f44da5-3e99-4cbb-80f9-d7d23d07b23e · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Alip: Adaptive language-image pre-training with synthetic cap- tion
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbd676ce-f04c-43ee-944b-1e300a7a59b0 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5cf8247-e0e0-49ea-a7ce-7f2cacd0fc0f · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a5d4ed-52e1-4694-9baa-893666e4f0ce · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7a435bc-b58a-4d30-bb5b-a618fd2e5103 · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Sigmoid loss for language image pre-training
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d5b9789-3e86-4abd-b014-b6b57742864a · outbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Dreamlip: Language- image pre-training with long captions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ffcd5aec-4fd6-4615-9460-aa3aa72f2de1 · inbound
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e9f23b-b96e-4e8a-a73a-bc3f5e4dfeb2 · inbound
Mining Contextualized Visual Associations from Images for Creativity Understanding CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f876f1b0-ae10-46cd-b18b-7130a3aece66 · inbound
MobileCLIP2: Improving Multi-Modal Reinforced Training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5172db-a689-4111-ae86-087c771bf7ab · inbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f7bbdb0-2b4d-4391-9dce-876277e74df8 · inbound
Let ViT Speak: Generative Language-Image Pre-training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e7a00b24-9bbb-4774-a8e9-96d0caa49fbf · inbound
Let ViT Speak: Generative Language-Image Pre-training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c35d4c3c-47fc-4f60-bc6e-2dff7871274c · inbound
Investigating Adversarial Robustness of Multi-modal Large Language Models CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66670d96-e570-4654-bdfa-0efce984068c · inbound
Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 46823c8a-a6ee-4e93-945b-adc5fa8d4ea2 · inbound
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.