Pith. sign in

Paper Citation Record · LEDGER

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models

As of 19 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.22431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22431 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:46:45.600246Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8687808d-6b90-469b-859c-9db72dba25ab · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.257026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.257026Z digest=sha256:68204229460a0229c93bf233438119249618cb623b4577fff71c44ae77ad92ed

Observation fdb97316-3ec5-4106-9d2a-27cd89f67a64 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.727507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.264665Z digest=sha256:3aa17494dd480564b145a212856b8715dacd48cdb02d3068786f48d4c51beee5

Observation e0755a2a-6891-4009-a98a-81960f9c5784 · outbound

This paper cites Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.697555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.275197Z digest=sha256:c8aa84e05d863f479a828e903c491dcaf97dc4106da8edf61fa08d9c0bf1e5b7

Observation 0313582b-5d45-43cb-9d05-8a60e610ddbd · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.677311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.285435Z digest=sha256:8b2b925f04686fdd97cc17532a6d4626971c82892784f558040b8a41e82ff0ff

Observation da9deaac-5bd3-4c4b-9d39-8fc13fc2bc11 · outbound

This paper cites Improving clip training with language rewrites.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improving clip training with language rewrites

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.656336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.293013Z digest=sha256:4989232637e61b23c4485aba324413bc9824222cfe148947cf53748e2c738414

Observation 8494b625-6a4d-4b55-ab54-5a42840d6363 · outbound

This paper cites Data fil- tering networks, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Data fil- tering networks, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.638350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.306448Z digest=sha256:483181b05b8742b527f4bc92d161c5b410fcce3e7fd7abda0f980cd1b53c15b2

Observation 039d2dd5-6b27-4bc1-963c-aab84da425aa · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.616900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.313820Z digest=sha256:c6901dae040b07a456066d9eb62de813a312ad644e2c9a391a1fc0878b39f46e

Observation 67bff9d6-c2f2-466e-b1f4-172f4223377b · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.588004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.325074Z digest=sha256:21683b80544b3f88d3536522ec59bd755403e7298f8aed50b4756ebdb91efb28

Observation 08175ae1-0b0d-48a8-a002-ca6b7aa2389d · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Datacomp: In search of the next generation of multimodal datasets, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.567545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.332237Z digest=sha256:c58f761d45fd3fecf8a34400a452743cdfeb4a871e8f8b5120a998ce21fe4222

Observation 7262f8be-d511-49d4-9f5f-7e83a661c015 · outbound

This paper cites Classification done right for vision-language pre- training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Classification done right for vision-language pre- training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.547919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.339790Z digest=sha256:08776607dd3a71c425fe63d194471ce5a2a47ca6cfdbe88bbf1af12c0711daa1

Observation 911d6cae-6135-43ee-abee-333269196a7d · outbound

This paper cites GPT-4o System Card.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.348101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.348101Z digest=sha256:8162de0e178e2c4f836408f7830a7b3ca6f7d6460331fec6c4483f952153b092

Observation f54851e5-6dbb-4eb5-a7ed-24e0d60f14d3 · outbound

This paper cites Open- clip, 2021.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Open- clip, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.528791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.362242Z digest=sha256:1a345b91c341c129a06aadd99a99b06383fc0ad19eedb14f8cb01f87e192865d

Observation f0fba9fa-74f1-4a7d-b16e-d53deb8a3785 · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions,.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Veclip: Improving clip training via visual-enriched captions,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.508293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.368122Z digest=sha256:654c9341002ae1d9b4bedfeb3d71ad151a166d1ed79189844a94ff877aac76b2

Observation 372feb51-e920-4792-a9ab-c9d3bc5b0ee6 · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.486667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.375675Z digest=sha256:adb236248e2112a8f8678077f76095d2e4620042e1068decc4f767b9e587974d

Observation 77ed29c0-809d-412b-8fff-de2ece218112 · outbound

This paper cites Grounded language-image pre-training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Grounded language-image pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.457246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.387890Z digest=sha256:723214e3053e259e6815f24c40f981e152926186ce23836ef6a89f40d9eb2cff

Observation 3fe468b9-21f8-477f-980b-b93aadf25fa3 · outbound

This paper cites Clipa-v2: Scaling clip training with 81.17, 11.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Clipa-v2: Scaling clip training with 81.17, 11

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.428986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.395048Z digest=sha256:948095b03dd644731515949e555c91597ac2d6c73f2ccebed5f540d42edcc9dc

Observation bca85c64-9114-4285-a012-c084a8a858ba · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.400816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.400816Z digest=sha256:5181f73a14c50986bbeadb94f1e53ee27e2dd57a0aced5d5ccd2030a227b6c1d

Observation 12443c6b-bad7-4196-ab35-6c714c0a0a5b · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improved baselines with visual instruction tuning, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.410742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.415558Z digest=sha256:68543dec780dc7b41f2cce0419b30a4e65a165a688418dfca82fbeb9df8f84fa

Observation cd5a54de-8e88-4196-9e92-ccfe55968d28 · outbound

This paper cites Visual instruction tuning, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Visual instruction tuning, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.420691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.420691Z digest=sha256:1552b65d45ceccd6498ef2ed0eacd91eb24a9c1d7e2da1b8c4bb5e1f1d45eda0

Observation 60a38ff6-e757-4a52-91d5-6debf7111c78 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.376042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.427199Z digest=sha256:688f7058224954c2d40e87f366f18b4f2cc3e93078f7a61301f9b6c2ae1266fa

Observation ab674b3e-c7a2-46c9-84cb-9726f7a5cfc4 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.432094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.432094Z digest=sha256:8c3afe4ed5d6b42df41f5cf2bdd03f332464717665a6fbbdd87f977a0616ebd9

Observation bae38e45-a8ab-44c0-866e-28cd93456e8a · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Slip: Self-supervision meets language-image pre- training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.331626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.440608Z digest=sha256:1c540441bce0682acf748e619c4a492ceb36e93e95166d73e5ff0a270894a54f

Observation 51154b03-b426-4fd2-bc89-38e3bcd0ef82 · outbound

This paper cites Improving multimodal datasets with image captioning.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improving multimodal datasets with image captioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.308562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.447199Z digest=sha256:89c03fb57d8ff973429056a768b0a48c636b4c317db916171031f63799437ff6

Observation 1a6f594c-487c-42a3-be44-5dbf127470de · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.282647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.453397Z digest=sha256:c0abe424baa0f3f02e41eef6a732fd0086573e0dd4a1287af1b8bfc271fbf0d7

Observation 21558c4b-2ba8-43ad-a5c6-19f6ce3b1c71 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Denseclip: Language-guided dense prediction with context- aware prompting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.459856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.459856Z digest=sha256:f3fcc4ef1d2c017a9111af6688e5ea8f78d8b501dc6cb7bfd2d481eb94c7156d

Observation d1b1a601-d1dd-4e87-a7f7-aaddb8730b77 · outbound

This paper cites Fusecap: Leveraging large language mod- els for enriched fused image captions.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Fusecap: Leveraging large language mod- els for enriched fused image captions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.237591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.465588Z digest=sha256:e42cde2dd4320ee5309f3c30c51f20e634f0f1dd36d22e66d4d1bea890391354

Observation 972fd9d1-1252-4f40-af1c-da919df8a99f · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.476634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.476634Z digest=sha256:a2b44881cf5a24b28247dd7436b091ce3df17b20b3dc77e4746e90b7e19801af

Observation 3d5e3a2d-5cc3-4ce9-839f-03e51b64573c · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.214046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.483145Z digest=sha256:684b7043251ee20a8ec365da0913d974df5816b71c534bbd9a3d690942c609a2

Observation 41e5be2f-a1ec-40fe-b8be-6310f94a4b85 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.488144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.488144Z digest=sha256:7f1639beb3d6e8117c3823c66660f358a4b2958047d2b176ceda7db5997ce5d0

Observation b64dc236-7bd4-4c47-b9e5-56161efecd1a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.494374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.494374Z digest=sha256:4e196113b4e6c93024389414304f1c869fe16f26f23054357edef9d005efe757

Observation 2e124188-8a4c-472c-8dab-0138ecb71694 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.499725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.499725Z digest=sha256:71e42a3dbec42011c3c3abf1142f9d48da7e70d4e23ed307274fe1d396ec4ec1

Observation 54248436-f71d-4947-a733-fae403771130 · outbound

This paper cites Locca: Vi- sual pretraining with location-aware captioners.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Locca: Vi- sual pretraining with location-aware captioners

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.184807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.505688Z digest=sha256:147d32f47bf2c52a5c2ae1623c20a289466bd2f53baa022430c08857e1eda8a0

Observation 1fbfab01-1955-4999-a42c-e42b7267011d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.512014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.512014Z digest=sha256:a7c79102581088fa681da547e97fed7a3621c92529a7a00fdd85ad0f3e8869fb

Observation 2335abde-285a-4cac-8447-b1a2c9cbfa25 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.517124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.517124Z digest=sha256:0d67914fea2668971e74fea399a026d33697da1d12f13a99a2ef19e461e10342

Observation 04d2f84a-85a6-43b4-b30b-3156a9177e25 · outbound

This paper cites Demystifying CLIP Data.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Demystifying CLIP Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.524128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.524128Z digest=sha256:503d810f7a02bb1036461521112264ff38fe53c7f84b0f940252e67f6b763896

Observation 8a8a5c90-6af3-4ba7-85d2-b63b449ccf57 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models, 2022.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Coca: Contrastive captioners are image-text foundation models, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.148854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.531912Z digest=sha256:e3fedd171c1ba159f1c4df74990bded7f2fba8aef1f57306a77a08dceaff9f8a

Observation 85b5c121-2c3c-4e77-9e74-d93c9127a5d4 · outbound

This paper cites CapsFusion: Rethinking Image-Text Data at Scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models CapsFusion: Rethinking Image-Text Data at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.543788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.543788Z digest=sha256:3a0b2a03030394bbf29e16e4529bde9791432e43838e3b11db098728547b8b01

Observation c059c51c-33df-4d19-b842-cd199730cd5a · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.116925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.550364Z digest=sha256:93817d75f042a8bf871df9997e5d7091b2698ae85636827e4c61fa77b9217a70

Observation 8b7f51ce-0d25-44d5-9ed3-2353ba1dacc2 · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In International Conference on Learning Repre- sentations, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models When and why vision- language models behave like bags-of-words, and what to do about it? In International Conference on Learning Repre- sentations, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.091776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.558671Z digest=sha256:bf00a0229b96a39e20d60f4b8b78c521fcaf257389b9463c7f61989b87e9c90f

Observation 9669d846-da58-4512-9db6-0f32adf4edfc · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.568273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.568273Z digest=sha256:0743d76259aa5c47ee63b07958bbe9d5cd251aa0f86f5fb9dffe68be366a38fe

Observation 0a7d25ac-ca9c-4a7c-8755-c435c58b6959 · outbound

This paper cites Sigmoid loss for language image pre-training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Sigmoid loss for language image pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.069538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.577179Z digest=sha256:74d838510de4f7b0e076958ba233a587dc67286e3a93d77162728634a44fa58c

Observation 8f631553-cbb2-41ca-b062-ad3de7d05214 · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Long-clip: Unlocking the long-text capability of clip

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.041210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.583471Z digest=sha256:aaa92243908095020640195672779e82bb969c5f02ee737535b3a8ac8c0763fc

Observation 4adfeb5e-0eca-42d0-8e52-8e9c2088bd60 · outbound

This paper cites Glipv2: Unifying localiza- tion and vision-language understanding.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Glipv2: Unifying localiza- tion and vision-language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.010744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.588353Z digest=sha256:bea1ea9698f691c2891fd4d9d2b5195e4c088f013d4629f69e2666532ba33510

Observation d794a3e4-5197-44fc-963e-5adfd3f1bf62 · outbound

This paper cites Training setup and dataset scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Training setup and dataset scale

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:45.988293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.593767Z digest=sha256:2696d10b49b085bac221a7d7abb05bdd149153021b62f62da4947d4fe3b0d7d8

Observation 72ead4ff-f661-4f49-ac37-0df59345b06e · outbound

This paper cites Examples We present some examples from the acquired dataset.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Examples We present some examples from the acquired dataset

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T11:46:45.961209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T11:46:45.600246Z digest=sha256:c8717ba8471dbc5e2f435f0fcfbc843eaff399a2575361cdf65d1336997989d8

Pith citing papers

No inbound Pith citation observations are available.