Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:46:45.600246Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.22431.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:46:45.600246Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8687808d-6b90-469b-859c-9db72dba25ab · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb97316-3ec5-4106-9d2a-27cd89f67a64 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e0755a2a-6891-4009-a98a-81960f9c5784 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0313582b-5d45-43cb-9d05-8a60e610ddbd · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation da9deaac-5bd3-4c4b-9d39-8fc13fc2bc11 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improving clip training with language rewrites
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8494b625-6a4d-4b55-ab54-5a42840d6363 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Data fil- tering networks, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 039d2dd5-6b27-4bc1-963c-aab84da425aa · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Eva: Exploring the limits of masked visual representa- tion learning at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67bff9d6-c2f2-466e-b1f4-172f4223377b · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08175ae1-0b0d-48a8-a002-ca6b7aa2389d · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Datacomp: In search of the next generation of multimodal datasets, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7262f8be-d511-49d4-9f5f-7e83a661c015 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Classification done right for vision-language pre- training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 911d6cae-6135-43ee-abee-333269196a7d · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models GPT-4o System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f54851e5-6dbb-4eb5-a7ed-24e0d60f14d3 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Open- clip, 2021
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f0fba9fa-74f1-4a7d-b16e-d53deb8a3785 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Veclip: Improving clip training via visual-enriched captions,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 372feb51-e920-4792-a9ab-c9d3bc5b0ee6 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 77ed29c0-809d-412b-8fff-de2ece218112 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Grounded language-image pre-training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3fe468b9-21f8-477f-980b-b93aadf25fa3 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Clipa-v2: Scaling clip training with 81.17, 11
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bca85c64-9114-4285-a012-c084a8a858ba · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12443c6b-bad7-4196-ab35-6c714c0a0a5b · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improved baselines with visual instruction tuning, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd5a54de-8e88-4196-9e92-ccfe55968d28 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Visual instruction tuning, 2023
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a38ff6-e757-4a52-91d5-6debf7111c78 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab674b3e-c7a2-46c9-84cb-9726f7a5cfc4 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mmbench: Is your multi-modal model an all-around player?, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae38e45-a8ab-44c0-866e-28cd93456e8a · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Slip: Self-supervision meets language-image pre- training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51154b03-b426-4fd2-bc89-38e3bcd0ef82 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improving multimodal datasets with image captioning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1a6f594c-487c-42a3-be44-5dbf127470de · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 21558c4b-2ba8-43ad-a5c6-19f6ce3b1c71 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Denseclip: Language-guided dense prediction with context- aware prompting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b1a601-d1dd-4e87-a7f7-aaddb8730b77 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Fusecap: Leveraging large language mod- els for enriched fused image captions
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 972fd9d1-1252-4f40-af1c-da919df8a99f · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5e3a2d-5cc3-4ce9-839f-03e51b64573c · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 41e5be2f-a1ec-40fe-b8be-6310f94a4b85 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64dc236-7bd4-4c47-b9e5-56161efecd1a · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e124188-8a4c-472c-8dab-0138ecb71694 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54248436-f71d-4947-a733-fae403771130 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Locca: Vi- sual pretraining with location-aware captioners
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fbfab01-1955-4999-a42c-e42b7267011d · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2335abde-285a-4cac-8447-b1a2c9cbfa25 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d2f84a-85a6-43b4-b30b-3156a9177e25 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Demystifying CLIP Data
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8a5c90-6af3-4ba7-85d2-b63b449ccf57 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Coca: Contrastive captioners are image-text foundation models, 2022
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85b5c121-2c3c-4e77-9e74-d93c9127a5d4 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models CapsFusion: Rethinking Image-Text Data at Scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c059c51c-33df-4d19-b842-cd199730cd5a · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b7f51ce-0d25-44d5-9ed3-2353ba1dacc2 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models When and why vision- language models behave like bags-of-words, and what to do about it? In International Conference on Learning Repre- sentations, 2023
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9669d846-da58-4512-9db6-0f32adf4edfc · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a7d25ac-ca9c-4a7c-8755-c435c58b6959 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Sigmoid loss for language image pre-training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f631553-cbb2-41ca-b062-ad3de7d05214 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Long-clip: Unlocking the long-text capability of clip
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4adfeb5e-0eca-42d0-8e52-8e9c2088bd60 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Glipv2: Unifying localiza- tion and vision-language understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d794a3e4-5197-44fc-963e-5adfd3f1bf62 · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Training setup and dataset scale
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 72ead4ff-f661-4f49-ac37-0df59345b06e · outbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Examples We present some examples from the acquired dataset
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.