Pith. sign in

Paper Citation Record · LEDGER

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 9 inbound Pith citation observations for arXiv:2411.16828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16828 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:57:12.054681Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:29:57.356586Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.707139Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87a9bce5-8af8-4d75-aece-df6652e9943c · outbound

This paper cites GPT-4 Technical Report.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.662098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.662098Z digest=sha256:7fa72cb9eb19e6c2c8f7c263c85477417233f6aaa5e896882edd03504b4ecbad

Observation 5d27400c-8ca4-4ae7-9221-91b6b7517cc0 · outbound

This paper cites nocaps: novel object caption- ing at scale.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions nocaps: novel object caption- ing at scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.393647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.668803Z digest=sha256:4fb64e82ce6a054771a3f296e1811547eed7c0e096bf6d0acc41833804da38b1

Observation 9c7aa08c-03a6-4128-be0d-7933828bdbe0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.674832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.674832Z digest=sha256:2e0178f6fe55475b4c7be656a0dcb8ed3f18e4fb88cbdd439abec0ab2609b2cc

Observation d0a64e86-c10a-46ae-b785-ac4b59bec140 · outbound

This paper cites Contrastive Localized Language-Image Pre-Training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Contrastive Localized Language-Image Pre-Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.682472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.682472Z digest=sha256:c1ed8b35c19e67b1743881295da4ee1eb7130cac05968761b0b5815da19efbab

Observation d865697b-4c7d-4f3d-8ded-4c99fe4736f3 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.689555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.689555Z digest=sha256:bcd6b069a94db1135cbdf9bf6d5e3cae1220c080c6594d9d157b51363381659d

Observation 2fa93cab-c879-401a-93bf-f6af44de8b9a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.695395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.695395Z digest=sha256:ef18b66b7b31e6b8775feacf5afa5c490f46c857360e98644bb438f0041bda37

Observation a6c63e58-115e-4a5c-9fcf-14d6cd4b683c · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.701555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.701555Z digest=sha256:e70b6db1bca452a73a9b6724e5a7a2c9badc5a7731bdb9fb2486b7ff66607391

Observation 31d3cbf0-4c91-43ab-b5ec-87d41022b000 · outbound

This paper cites Uniter: Universal image-text representation learning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Uniter: Universal image-text representation learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.710436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.710436Z digest=sha256:8ee306913c4ecafec2cc925e7917b1d4b90e0142281dfd55d9c6230fb71da46c

Observation 856507f4-b276-45cd-b87d-591fca256d3d · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.721096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.721096Z digest=sha256:c49b68cbddf65a3f648ec621fc214e49b973990ffde6d1dc63e2eb902ad621f8

Observation 19eac34b-3930-450a-a17e-cc60b3b72e55 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.727721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.727721Z digest=sha256:2583aca6c34f79591b063df2be2ca734705cebfa706ef9c0b135757655eea389

Observation f8190cd5-b06b-4614-854d-11611a82aa61 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.734320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.734320Z digest=sha256:b39fd2300da210e4b235a5723f485ba25d3ff8b9ac443519465979efaad16c39

Observation f2ebdab8-4d15-4188-a959-e3da8492db67 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.743434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.743434Z digest=sha256:ff5bca52dc30e3bcbb50e3ee128ba9d34f9a1897fbe09c117e0471d970d50587

Observation b9f5d061-4bfd-406c-acde-9f0c709a33cf · outbound

This paper cites The Llama 3 Herd of Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.751560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.751560Z digest=sha256:d758fcdbec393b692c7fda34916e355d39db35266b92417f2bb0823086084615

Observation 82a20a60-25b5-4c7f-a116-bb3244b52271 · outbound

This paper cites Improving clip training with language rewrites.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improving clip training with language rewrites

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.320690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.758151Z digest=sha256:e511f91faea4143ab0ea6710266f42a4a010d6e1933380404f135bff9e4e6afb

Observation 23714bab-d68e-4368-b4ad-f8eb6dbd2213 · outbound

This paper cites Data Filtering Networks.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Data Filtering Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.763989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.763989Z digest=sha256:945e627369d9dab685bff26db72a4d9cbd2d5fd0cc79f4be2f0e965c64e89c30

Observation d164bd10-f6a1-4f7e-a598-75097867be19 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.771942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.771942Z digest=sha256:5cee23e56f2e69b676ea7a447825aa5a2d67edd8313a73041982e5ac5b612eda

Observation 675b36ef-7343-423b-a2b9-774c7b581dac · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Dat- acomp: In search of the next generation of multimodal datasets

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.277650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.780479Z digest=sha256:343b41405725edb42d0ff77ad9ba1d516708f769cdc71602e755083b32ccc314

Observation c0e496eb-6e35-46fc-8f0d-99aa2501f4af · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.786071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.786071Z digest=sha256:ba9d84f3f461164b5abd36fff496edd6feba020377a2ad7672fb8369786184b6

Observation d0e9be35-437e-48ee-be88-39a63aab6713 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.247902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.807564Z digest=sha256:d9f60ee8ace085fe5c1704f312aa2d2d7b96244df78a70116917131aba81554b

Observation 7e1d1dc0-b2cb-4289-961c-7305b5a4e58e · outbound

This paper cites Open- clip.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Open- clip

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.217337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.817392Z digest=sha256:1b80a7fabcf16ff7be01f66ff4598b573a88f647d065c39a8fa854825eb1857e

Observation 4e663a27-f256-4864-8de8-a8110ad34162 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.823008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.823008Z digest=sha256:a4ded432a13425139461e70b31f70c825fbd142342046f7debd44bf0ed949caa

Observation e77d857e-3b87-4a6d-9e35-25bcd6e60d66 · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.829469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.829469Z digest=sha256:16720292a5019cc9fbd5b6c8d15c9e774221404cf4ae8242af3b8c2931f3953e

Observation 2fafa437-3385-432e-ad67-33adf7348b28 · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Veclip: Improving clip training via visual-enriched captions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.164438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.835490Z digest=sha256:489580bb31a835081ccdbd64f01ccf473a5f3c3d03c89d47bc460378d4a196b8

Observation 9299736a-f2b8-4ea5-9d26-fa4dbc5a0f44 · outbound

This paper cites Modeling Caption Diversity in Contrastive Vision-Language Pretraining.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Modeling Caption Diversity in Contrastive Vision-Language Pretraining

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.841835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.841835Z digest=sha256:bf1d03cb16214f377ddc7a6c599a42a8029cbe243ec17c5d5dba8940b4b01a0d

Observation 94c672f7-5eae-4a63-a2d5-3b53ef73108d · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.849373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.849373Z digest=sha256:13665d439e3066dd8810f3d261ed74ba29e4f773ec84ad0727d743eb0986ea76

Observation c46fb3ea-2ed2-4ca7-a8c7-00abf378e968 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.857195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.857195Z digest=sha256:843d6fbfb2f759db5a43b54201723013240f067b4801f17abf3ff7c300d16535

Observation fed288c3-f5a6-467a-aa17-970573b7dd0f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.863570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.863570Z digest=sha256:811609b2734bfcc09e019c3ac054bdd86595af1d38a7e356e69f0007ddb30345

Observation d1171e4b-2022-4e26-9bcf-245de887639f · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.870361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.870361Z digest=sha256:f66ad56e0859b311b2697458a0c9b7c246f2253256bfaa0c8f696face5a51bfe

Observation 4886e482-f71b-40e5-baa4-cff250d1c4af · outbound

This paper cites CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.877950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.877950Z digest=sha256:cbd1b4277fc8ec59b1987830af20109fb91163d9a1951b0858ee9bcc9bec8ca7

Observation 40b79ca7-7d68-4d30-ae59-f2fc638f8802 · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions What If We Recaption Billions of Web Images with LLaMA-3?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.884915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.884915Z digest=sha256:c7a312457cf4741bfa4a0a5d2843300330713e19abb618a48534d06a05078568

Observation 118f6ee0-5eb3-442b-834f-a3b6a43de51f · outbound

This paper cites An inverse scal- ing law for clip training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions An inverse scal- ing law for clip training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.096232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.890416Z digest=sha256:7d23fedcd44458f4a3c94bf06c6971349fdee8fa6908a111ebe6118574618d29

Observation 3d6e3517-fe59-4b9f-8675-3d1cd0181484 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Evaluating object hallucination in large vision-language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.895324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.895324Z digest=sha256:b2e1ad29679819a1118c4908e1f91937d6475e130bb8af31d33eed31d520a41c

Observation 2e13e9a0-7ba8-4341-8a23-3029afa06591 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.900633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.900633Z digest=sha256:9fd5b586b569daa6bb654413f816ad536d2493efdbf95ac495051c97649a005a

Observation b20fcd7b-2378-4c47-be2c-be2e2a52e369 · outbound

This paper cites Microsoft coco: Common objects in context.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Microsoft coco: Common objects in context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.912062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.912062Z digest=sha256:9bf148e71518de6b5652517e6d91dc56b5bafb50ae101683d5e81badf3cfb2b0

Observation 4c7bcbe9-72e3-4cc9-a67d-2f9137696731 · outbound

This paper cites Improved baselines with visual instruction tuning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.027445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.917940Z digest=sha256:6036ac576dc49797e09a2146af837f0ea4d9174b7c3e1cc09732340dbf412d90

Observation 974a6fed-75b7-42fe-a859-4a80f5947e52 · outbound

This paper cites Visual instruction tuning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.923586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.923586Z digest=sha256:a34a0852c9c54aec77c34c6e60f12e68b6857d376742f9fe9f93d88a63424c20

Observation 2def70fc-4104-4b95-9f14-89fd72aaa89d · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MLLMs-Augmented Visual-Language Representation Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.929501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.929501Z digest=sha256:38354b3030637e24d8e5bc3dcb89c0aaa7042e192efb126200c7f42ea6b0ef13

Observation 79aad07b-cc57-4341-828f-90a2f2aff7a7 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.937991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.937991Z digest=sha256:4f63cd66ca5760be58d0c93226ccde4fb9a22a2c9a8a6b95a55d04bbd3c9503f

Observation b816f311-eb6c-47b7-8a25-e3285280ef85 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.961870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.942742Z digest=sha256:59599c525bc58b9c24aa55165df6cc43644d7edd4ea14dc4f6283cf7757827a3

Observation 1f25dc0d-99d3-4021-8a8e-2ab38f1d389f · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Introducing meta llama 3: The most capable openly available llm to date

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.940951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.950195Z digest=sha256:06df4345e8b7cd43e7b7e17bdaf4859676ecd4eb21d8a1e31fa090af94ecde5f

Observation 95029b9b-1135-47f5-9d57-a8d123894450 · outbound

This paper cites Improving multimodal datasets with image captioning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improving multimodal datasets with image captioning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.918110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.957998Z digest=sha256:6e017ae40ffcb34eb9cb15eb628f20713489955615b5e07c92cfd0ac72e93c10

Observation 864f0c83-36d8-4123-8902-435d856882bb · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Representation Learning with Contrastive Predictive Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.963589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.963589Z digest=sha256:a12b9cd780b6319d5e44accf37c0e8ad501f2c1b6139ded7a0215b88ee4d6b56

Observation 9d418dd9-57d1-4748-8765-a003e296d0f6 · outbound

This paper cites Introducing chatgpt.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Introducing chatgpt

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.889022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.969254Z digest=sha256:851488d4b2c34fbaf4b9eb23ca80e19c6ecc580f3a6073ecaa98a035c6980d3b

Observation 7bc454e2-6576-4262-a68c-d1ddbcdcf4b4 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.974372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.974372Z digest=sha256:d2c0adb0ce2be25278190794ff7738bb06c6623de1623558ccafa3a5dbe26b19

Observation cdf99a94-d00c-45e2-846b-2441208eb8da · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.850186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.979462Z digest=sha256:ac139edf00c923010de520e78e5146ec9662a42f407564cf806ad4f176a29584

Observation d2d1b4a7-cc30-46b9-b26b-46a15f778456 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.828502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.987926Z digest=sha256:4be7bd096853e8f74b1f5f231ab4a5634abce7d7480d1639ea4bd399e53dd08c

Observation 717c274e-ca7c-44aa-9a32-90881f0de1df · outbound

This paper cites Towards vqa models that can read.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Towards vqa models that can read

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.805935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:11.995298Z digest=sha256:a337b51739446bc9d872f50076522c8b4f5993108e98ddb3bee1b43171c8d127

Observation 35637afb-a9e7-457d-99da-76ec8784bc20 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.778594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:12.003199Z digest=sha256:6035ddc56f6326cf992a43cbda65da17224d0030f672b67f4c29b370390e0688

Observation d8bc428a-7b52-45df-a683-b0a6d8981a14 · outbound

This paper cites MOFI: Learning Image Representations from Noisy Entity Annotated Images.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MOFI: Learning Image Representations from Noisy Entity Annotated Images

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.015082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.015082Z digest=sha256:7b265760b2178b3a66231b6f1415bf82b0462ba574ccd8d4cc99cced9721533d

Observation b9f44da5-3e99-4cbb-80f9-d7d23d07b23e · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic cap- tion.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Alip: Adaptive language-image pre-training with synthetic cap- tion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.021017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.021017Z digest=sha256:8587e3ae48a05f5419c634fa36d0cb249ddfd90f9f05d7aff4080ee13f576feb

Observation dbd676ce-f04c-43ee-944b-1e300a7a59b0 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.027307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.027307Z digest=sha256:863dbdf5f1dbe4c7af738918e22a66e7b0b4b32af191f78b3ef2d3797194e949

Observation f5cf8247-e0e0-49ea-a7ce-7f2cacd0fc0f · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.033282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.033282Z digest=sha256:7b462cc2d46e4e319344053132d006868a1b4fc212e23a750c64aa60cd514850

Observation 86a5d4ed-52e1-4694-9baa-893666e4f0ce · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.744561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:12.040523Z digest=sha256:046b6904c953a4592defd47b3534162bd64ff39b97e22a1806593ee6068ca81f

Observation b7a435bc-b58a-4d30-bb5b-a618fd2e5103 · outbound

This paper cites Sigmoid loss for language image pre-training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Sigmoid loss for language image pre-training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.724421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:12.048353Z digest=sha256:dd88f11ba4cb6ee9f9dcbe62e97f22ffc00d52d3b253e850797fa2d052928056

Observation 0d5b9789-3e86-4abd-b014-b6b57742864a · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Dreamlip: Language- image pre-training with long captions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.695816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T12:57:12.054681Z digest=sha256:ffac7513b40edeba74a434b1bcfc945cdfc87045b68861baabdcef0a13a1b2d9

Pith citing papers

Observation ffcd5aec-4fd6-4615-9460-aa3aa72f2de1 · inbound

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning cites this paper.

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:29:57.356586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:29:57.356586Z digest=sha256:577cf141171b3f96e2c96edf8b6d9353da105b884250ad9b618d8075ba12ace7

Observation 94e9f23b-b96e-4e8a-a73a-bc3f5e4dfeb2 · inbound

Mining Contextualized Visual Associations from Images for Creativity Understanding cites this paper.

Mining Contextualized Visual Associations from Images for Creativity Understanding CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:09:27.053527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:09:27.053527Z digest=sha256:5a8134dd20769e5378444b61c35561dd207c58eb3a4ccf9e4a6f789e9f3399db

Observation f876f1b0-ae10-46cd-b18b-7130a3aece66 · inbound

MobileCLIP2: Improving Multi-Modal Reinforced Training cites this paper.

MobileCLIP2: Improving Multi-Modal Reinforced Training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:19.577024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:19.577024Z digest=sha256:835ffea1aaf6c88f4b7ad93e8197c2b40144c03a749173de11676f4b637e19cc

Observation bd5172db-a689-4111-ae86-087c771bf7ab · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.606923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.606923Z digest=sha256:13413a8d8df4cec1b89ac4944de7a005b92ceb129cfd240a618274cabecc6c8e

Observation 7f7bbdb0-2b4d-4391-9dce-876277e74df8 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:07.255186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:409483ad1d4e2ee974f8d5d801edbb6f45237a7971d1d11e4e76d899fb8c9d02

Observation e7a00b24-9bbb-4774-a8e9-96d0caa49fbf · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:28.764723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:60d3e953e32bd00c8f0f0f6301a19c15388c8b61466fd46bd1a855870f11dc82

Observation c35d4c3c-47fc-4f60-bc6e-2dff7871274c · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.577528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:a7151c074ed9df87ce79af203f504ebd73b5441212263cef4fa262b326131351

Observation 66670d96-e570-4654-bdfa-0efce984068c · inbound

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models cites this paper.

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.830996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:04:30.654255Z digest=sha256:57a14a899dde16c59303afe21e51c5ccd79b3ca8ac234871fbafbd8c26431ecf

Observation 46823c8a-a6ee-4e93-945b-adc5fa8d4ea2 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.708849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:eadd1c5420c89fe49a8ca03bda9ea7903506b802ece3b6506762d9a46ce07455