Pith. sign in

Paper Citation Record · LEDGER

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 9 inbound Pith citation observations for arXiv:2411.16828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16828 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:57:12.054681Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:29:57.356586Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.707139Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87a9bce5-8af8-4d75-aece-df6652e9943c · outbound

This paper cites GPT-4 Technical Report.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.662098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.662098Z digest=sha256:f92747d7c48aafc4c0702adc14321ad74a7f24072c018c421cde27f603fe6c0d

Observation 5d27400c-8ca4-4ae7-9221-91b6b7517cc0 · outbound

This paper cites nocaps: novel object caption- ing at scale.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions nocaps: novel object caption- ing at scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.393647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.668803Z digest=sha256:4e9e7dd724c49fdc41f918576fd8dc7fec159e40a33f4e4db40837f00f7ddfe2

Observation 9c7aa08c-03a6-4128-be0d-7933828bdbe0 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.674832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.674832Z digest=sha256:ef207df44834ed243cfef1c7bbd4e5f8b408ebdc54910a3de22a8efe90b44a1e

Observation d0a64e86-c10a-46ae-b785-ac4b59bec140 · outbound

This paper cites Contrastive Localized Language-Image Pre-Training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Contrastive Localized Language-Image Pre-Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.682472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.682472Z digest=sha256:604b04ca8d5956c57e1d64865ac786c45b3f9e1ecf3bd880cf30f502f2616c6a

Observation d865697b-4c7d-4f3d-8ded-4c99fe4736f3 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.689555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.689555Z digest=sha256:e5483655291a36ff8b4464607795873eeec1414c12129247677f1a3a18cd07bf

Observation 2fa93cab-c879-401a-93bf-f6af44de8b9a · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.695395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.695395Z digest=sha256:752e3685140dd8fdd235d92684e2bbb7df67571d958f8601ad2b1d01ea32e186

Observation a6c63e58-115e-4a5c-9fcf-14d6cd4b683c · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.701555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.701555Z digest=sha256:20848a685f856a8090848d02d722d3046c45288ba9559bea36e88c301a3e0819

Observation 31d3cbf0-4c91-43ab-b5ec-87d41022b000 · outbound

This paper cites Uniter: Universal image-text representation learning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Uniter: Universal image-text representation learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.710436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.710436Z digest=sha256:fb3c83b7b6a5b1437a1b79d0b1ba1fc4abe3603a30084c6ae0f7de6548a8a6ae

Observation 856507f4-b276-45cd-b87d-591fca256d3d · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.721096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.721096Z digest=sha256:3a78a6ad6ef18eecbe3590f51c4afe57b027bfaec0fc7e171c829af0adbc58b6

Observation 19eac34b-3930-450a-a17e-cc60b3b72e55 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.727721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.727721Z digest=sha256:1101ebb322ec86b6e3729cf6f0db37c70a0d0b1c097482e6cfda45773fd776e3

Observation f8190cd5-b06b-4614-854d-11611a82aa61 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.734320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.734320Z digest=sha256:ca172e9562ce9dfcf26bebe8de2ade7065f0cce597e2cbaa4df44ff90fc60f6d

Observation f2ebdab8-4d15-4188-a959-e3da8492db67 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.743434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.743434Z digest=sha256:826f11596df9f942c35146077b702b7f4d2d992cff38ffbd665dc4ca546065bf

Observation b9f5d061-4bfd-406c-acde-9f0c709a33cf · outbound

This paper cites The Llama 3 Herd of Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.751560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.751560Z digest=sha256:d319a14c457fe1c9c3105c1f1f7d0cc7007aa8d5aa68d4fd4e2d5ca0c24b079d

Observation 82a20a60-25b5-4c7f-a116-bb3244b52271 · outbound

This paper cites Improving clip training with language rewrites.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improving clip training with language rewrites

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.320690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.758151Z digest=sha256:028ac616f1b1d883be32c7b4259edc42ae6839dde042fe1968e61a508029b207

Observation 23714bab-d68e-4368-b4ad-f8eb6dbd2213 · outbound

This paper cites Data Filtering Networks.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Data Filtering Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.763989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.763989Z digest=sha256:b1dfa05f6cbd097fc756e007ff134ce48fba3dd404323f8a367220ce7605bb7b

Observation d164bd10-f6a1-4f7e-a598-75097867be19 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.771942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.771942Z digest=sha256:4ca3c010da4d727a1a905622726e2ab96555ad7a446e9401aa546792c9ae0baa

Observation 675b36ef-7343-423b-a2b9-774c7b581dac · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Dat- acomp: In search of the next generation of multimodal datasets

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.277650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.780479Z digest=sha256:d484f234c20c420a91e4fe761eda2dc6ab394a3bd87edd07857a131f563cc175

Observation c0e496eb-6e35-46fc-8f0d-99aa2501f4af · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.786071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.786071Z digest=sha256:bb56978e68fc4c331999f3b5caf203513e7ea7e87426b9482572b7c121ab4d82

Observation d0e9be35-437e-48ee-be88-39a63aab6713 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.247902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.807564Z digest=sha256:5562929c2a58d29c14fb76703f9ad8d00488aa934fe5614725d6d7acd538e8db

Observation 7e1d1dc0-b2cb-4289-961c-7305b5a4e58e · outbound

This paper cites Open- clip.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Open- clip

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.217337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.817392Z digest=sha256:5e5f35042e2259259ded139b87e7cbb3b8dfcdb717b311eec3a413fff1b2b429

Observation 4e663a27-f256-4864-8de8-a8110ad34162 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.823008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.823008Z digest=sha256:10ec7f18ab1741ceeac9eadb231e682e435de608f4a18512ce4611d83d48b3fe

Observation e77d857e-3b87-4a6d-9e35-25bcd6e60d66 · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.829469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.829469Z digest=sha256:0eed52b2f9228bd05d09b653a965705cbadad6a4b4965e099bb1a24bc369c526

Observation 2fafa437-3385-432e-ad67-33adf7348b28 · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Veclip: Improving clip training via visual-enriched captions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.164438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.835490Z digest=sha256:89f573dafd51947a5bf2e8aa5d0bbc286f7c94af51e372e14878f362cb004c95

Observation 9299736a-f2b8-4ea5-9d26-fa4dbc5a0f44 · outbound

This paper cites Modeling Caption Diversity in Contrastive Vision-Language Pretraining.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Modeling Caption Diversity in Contrastive Vision-Language Pretraining

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.841835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.841835Z digest=sha256:a6ad1a512e985a83d20daf465e83d3b0b3579a3796daa9c8fcc28f9c03c5e472

Observation 94c672f7-5eae-4a63-a2d5-3b53ef73108d · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.849373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.849373Z digest=sha256:2eb9aa428f66313e0d500149350da4e9944c91ccb11df65d63338e3772fe2638

Observation c46fb3ea-2ed2-4ca7-a8c7-00abf378e968 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.857195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.857195Z digest=sha256:739405ce13e6b19aa8958b2b29e9c7c04e7007902b6fb1d6d21a9a056fc5faf3

Observation fed288c3-f5a6-467a-aa17-970573b7dd0f · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.863570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.863570Z digest=sha256:994454250f45dfc220e11936e6c09005055a989ffe0e5b2e6429b45e352677f3

Observation d1171e4b-2022-4e26-9bcf-245de887639f · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.870361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.870361Z digest=sha256:3752a631603b9db9120abded5f37dca95a374e39ee6054bd688bd4cf23e6817e

Observation 4886e482-f71b-40e5-baa4-cff250d1c4af · outbound

This paper cites CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.877950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.877950Z digest=sha256:29b2eedb64e772d4e02a022cde67caaf65ff9b0a7f0b8d0f15989d4f59615c3b

Observation 40b79ca7-7d68-4d30-ae59-f2fc638f8802 · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions What If We Recaption Billions of Web Images with LLaMA-3?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.884915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.884915Z digest=sha256:20564a46961c0f060442588901238482e37095a3986379d12c84f157584d3d92

Observation 118f6ee0-5eb3-442b-834f-a3b6a43de51f · outbound

This paper cites An inverse scal- ing law for clip training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions An inverse scal- ing law for clip training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.096232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.890416Z digest=sha256:e449b0fd8c875565f56584fbcc724cb5ab35d7147d67a1c5351ec6ec4f43146a

Observation 3d6e3517-fe59-4b9f-8675-3d1cd0181484 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Evaluating object hallucination in large vision-language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.895324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.895324Z digest=sha256:8483db724f6d1d86fd1a551a941d49c3fe5a2972eb9cec19fe567f13eec1cd5b

Observation 2e13e9a0-7ba8-4341-8a23-3029afa06591 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.900633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.900633Z digest=sha256:0b3961d1ab386fccb6f77188b5924f779767c00f67de2c02ca66468a7666b745

Observation b20fcd7b-2378-4c47-be2c-be2e2a52e369 · outbound

This paper cites Microsoft coco: Common objects in context.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Microsoft coco: Common objects in context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.912062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.912062Z digest=sha256:b14a8ad649120be23a08de6bced021a55601275be04f33b0490837f726e226b8

Observation 4c7bcbe9-72e3-4cc9-a67d-2f9137696731 · outbound

This paper cites Improved baselines with visual instruction tuning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:13.027445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.917940Z digest=sha256:d17806cbabe6d4e560ce432dca1c85d321821abcbcefeebd3f5b48277b169a18

Observation 974a6fed-75b7-42fe-a859-4a80f5947e52 · outbound

This paper cites Visual instruction tuning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.923586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.923586Z digest=sha256:e5497e75bf4ce11543a199894a488228bba1bc4982ef9d3286e6177397b49853

Observation 2def70fc-4104-4b95-9f14-89fd72aaa89d · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MLLMs-Augmented Visual-Language Representation Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.929501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.929501Z digest=sha256:43c1cb77b65f8cd06bde5c8a34d2f050dfb9c03955dfb5ca6896d64ec713643a

Observation 79aad07b-cc57-4341-828f-90a2f2aff7a7 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.937991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.937991Z digest=sha256:6d3557f6568ad749ee907ccccf72895399fbd44097e7a3bc6d0406fd95ca62db

Observation b816f311-eb6c-47b7-8a25-e3285280ef85 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.961870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.942742Z digest=sha256:88c71efda52fa859ec37e3e2b75312ef8f12b385aa26d709b38c8f7bb46d5d66

Observation 1f25dc0d-99d3-4021-8a8e-2ab38f1d389f · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Introducing meta llama 3: The most capable openly available llm to date

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.940951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.950195Z digest=sha256:35e3e3003d8ba5268220293fae937b85dbc46fad2f3d87c0aa59e3706e270050

Observation 95029b9b-1135-47f5-9d57-a8d123894450 · outbound

This paper cites Improving multimodal datasets with image captioning.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Improving multimodal datasets with image captioning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.918110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.957998Z digest=sha256:e8a5f15466f2c3b1ccd2fc92258631139653730fb47de24acdc3198cfe9003e0

Observation 864f0c83-36d8-4123-8902-435d856882bb · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Representation Learning with Contrastive Predictive Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.963589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.963589Z digest=sha256:50bfffb04b96fcbd71fd5d0d00beed7c3690edcc39548e2f98e35f4558a30cf9

Observation 9d418dd9-57d1-4748-8765-a003e296d0f6 · outbound

This paper cites Introducing chatgpt.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Introducing chatgpt

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.889022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.969254Z digest=sha256:4d486a2bf3a6962ab28777205337c9757de2ac46d68ebcea96889a5c8d5165bf

Observation 7bc454e2-6576-4262-a68c-d1ddbcdcf4b4 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.974372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.974372Z digest=sha256:23147796f4812aefbd59ab9ec9facaa28e13e470dbd2d0bcd1177a7b9a31f43a

Observation cdf99a94-d00c-45e2-846b-2441208eb8da · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.850186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.979462Z digest=sha256:539c8fe5ce3176b95c9f0c1fd8f42d204b3733d7cd99cfd0dd338c337698908e

Observation d2d1b4a7-cc30-46b9-b26b-46a15f778456 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.828502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.987926Z digest=sha256:759f06362282ff9ed12b7e4d2da9e03e80d0bc478c0070dfd3b6c2012fc68b75

Observation 717c274e-ca7c-44aa-9a32-90881f0de1df · outbound

This paper cites Towards vqa models that can read.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Towards vqa models that can read

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.805935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:11.995298Z digest=sha256:e95bc89549215f8d914437e40996966c4b0a32ed17452ece076e2890ba5d8aef

Observation 35637afb-a9e7-457d-99da-76ec8784bc20 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.778594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:12.003199Z digest=sha256:7c08dba7b926dc49bb9b8103ee9adbabb81a4e31c907ca6d5639c909e3fc68ac

Observation d8bc428a-7b52-45df-a683-b0a6d8981a14 · outbound

This paper cites MOFI: Learning Image Representations from Noisy Entity Annotated Images.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions MOFI: Learning Image Representations from Noisy Entity Annotated Images

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.015082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.015082Z digest=sha256:7a9a5f337af4c921f3134ff3e216f92f88962fe93cda140e8174a048671dfdd6

Observation b9f44da5-3e99-4cbb-80f9-d7d23d07b23e · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic cap- tion.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Alip: Adaptive language-image pre-training with synthetic cap- tion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.021017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.021017Z digest=sha256:72cfc43271f45da32858a5aa50a1b1af0a4ea9351e8153029e0907ccd1b871f6

Observation dbd676ce-f04c-43ee-944b-1e300a7a59b0 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.027307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.027307Z digest=sha256:aaaf870854cd66b67f6a5058fbd2456ac7d29d5cabdaaaad757c090853f9a8a0

Observation f5cf8247-e0e0-49ea-a7ce-7f2cacd0fc0f · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:12.033282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:12.033282Z digest=sha256:f4cfc80b15ca476fa9a251eb095cd218b19325b2ce2baed9d57e09348fd2b980

Observation 86a5d4ed-52e1-4694-9baa-893666e4f0ce · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.744561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:12.040523Z digest=sha256:95c856fe672dd9ade431942e26f1442f944326318ef47995e317954de8ee4638

Observation b7a435bc-b58a-4d30-bb5b-a618fd2e5103 · outbound

This paper cites Sigmoid loss for language image pre-training.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Sigmoid loss for language image pre-training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.724421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:12.048353Z digest=sha256:748826ec8dbffb4fa0871656fa0aa28ac367ad750e51f070e5b77730262dcba6

Observation 0d5b9789-3e86-4abd-b014-b6b57742864a · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions Dreamlip: Language- image pre-training with long captions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:57:12.695816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:57:12.054681Z digest=sha256:7873062d915c519c897604c8a113c38edf5d5e1c91a5990e25c9cadcfbef3c6d

Pith citing papers

Observation ffcd5aec-4fd6-4615-9460-aa3aa72f2de1 · inbound

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning cites this paper.

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:29:57.356586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:29:57.356586Z digest=sha256:1eb5acc43ebcbd9c9a40db50ce4f59f62328311da746e41980a5651457902115

Observation 94e9f23b-b96e-4e8a-a73a-bc3f5e4dfeb2 · inbound

Mining Contextualized Visual Associations from Images for Creativity Understanding cites this paper.

Mining Contextualized Visual Associations from Images for Creativity Understanding CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:09:27.053527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:09:27.053527Z digest=sha256:dc935cc10b945c8891ac93f6a33dcccebf7fbc1bcde9f24547c9f96697120edc

Observation f876f1b0-ae10-46cd-b18b-7130a3aece66 · inbound

MobileCLIP2: Improving Multi-Modal Reinforced Training cites this paper.

MobileCLIP2: Improving Multi-Modal Reinforced Training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:19.577024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:19.577024Z digest=sha256:2f9e8c3812f4b00e94fdb0a1ae38a14fa319d480dff52addd7d1d09bad4ce523

Observation bd5172db-a689-4111-ae86-087c771bf7ab · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.606923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.606923Z digest=sha256:81e3b386379f98231cf67af2402d11cbfef154ce694a9c846780f311966688ea

Observation 7f7bbdb0-2b4d-4391-9dce-876277e74df8 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:07.255186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:faad3a5b1f2bb5906eaa9bc1d472347af8ac97061e8289d4f27387607c35d22a

Observation e7a00b24-9bbb-4774-a8e9-96d0caa49fbf · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:28.764723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:2ab03cebd6f83b8db34f5520f2365c88feea6da1e3461cf927ddb2d700a9bc86

Observation c35d4c3c-47fc-4f60-bc6e-2dff7871274c · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.577528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:18ab5c69c53846777ec663f19f26083833df1c33ca3e39ca7ecd2c767ecacb81

Observation 66670d96-e570-4654-bdfa-0efce984068c · inbound

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models cites this paper.

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.830996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T11:04:30.654255Z digest=sha256:f0501397f189b6c8fc6cddb8694b4c809bf451ced3d55337b15f5e599dae1d09

Observation 46823c8a-a6ee-4e93-945b-adc5fa8d4ea2 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.708849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:42a6ae04bec30386a7b60d5b6e7cf02357ff9b8a5a1d926e05722325df9e0e55