Pith. sign in

Paper Citation Record · LEDGER

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

As of 13 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2411.11927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11927 v3

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:19.066615Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:20:50.578691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fc2996f-d8e4-4fdc-8e1c-b26acbc91c16 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.787439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.787439Z digest=sha256:8656620b0666aeb421f5f8fa0144e4c34cc12100bff46c0520d72ba007b4fd39

Observation 771fe683-689f-441d-b6e9-682a5f800c57 · outbound

This paper cites GPT-4 Technical Report.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.792629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.792629Z digest=sha256:25f95ca2857ce187a3173e742fac12c90e92f56f4d70d0250f573d280cc8a258

Observation 660777d5-cb9c-4b98-b151-3a197f179d4e · outbound

This paper cites Qwen Technical Report.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.797615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.797615Z digest=sha256:383b5f85bd68fafb6b5a16d37e3276efee353de6d90b83720943847e92a891d6

Observation ffdf892f-bd6e-4a73-abd7-c55e88b008fd · outbound

This paper cites Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.802194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.802194Z digest=sha256:8d4f5fab54cb9c30b0815af5e5b6f16f63b4b26d2941a5da02b00b0202f51e39

Observation 4c869d3f-597c-4e61-92cc-3786a74a9fb0 · outbound

This paper cites Ar- trackv2: Prompting autoregressive tracker where to look and how to describe.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Ar- trackv2: Prompting autoregressive tracker where to look and how to describe

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.991922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.806733Z digest=sha256:71f33fb1dde13946efba23969401d8b6a7e5a15515cb959b3abef52124c039d7

Observation c4c56f93-3de8-484b-a1ad-86a340e98125 · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.810790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.810790Z digest=sha256:a0256a07b47a18256657571c6e0d17d6db10456d5fb8fa62b6bc6497217c4bf0

Observation 5e25dad2-cdbc-4c67-a9af-08fa36885ddf · outbound

This paper cites Language models are few-shot learners.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Language models are few-shot learners

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.981028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.815209Z digest=sha256:c186188d50f0cfee941fa19559f738edf24fb705b32faf9d086a165e4112c05b

Observation 0c0df006-4258-4b48-ba93-2aa5bcbf47b9 · outbound

This paper cites Cross-lingual and multilingual clip.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Cross-lingual and multilingual clip

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.968846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.819267Z digest=sha256:fb0be4f287c409f4fdac9396d9cf6c6e9142b3ffae50618fdc60262ef3fd5547

Observation f2176700-0419-4c48-b9f3-f506a023749c · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.823234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.823234Z digest=sha256:7254090449e4064d54b957222ab81d0715bf8185b1bc53061173fb1128bc4e3f

Observation 99a2709e-9a1a-4d99-890b-3cf78d420f12 · outbound

This paper cites Pali: A jointly- scaled multilingual language-image model.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Pali: A jointly- scaled multilingual language-image model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.954196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.827882Z digest=sha256:a07169e606ed9072db1ac1b6fd3ac7cbba6ae8f68f9a7486ca4e78f419b80da5

Observation d1b03554-275a-471a-bd59-81be51e97053 · outbound

This paper cites Altclip: Altering the language encoder in clip for extended language capabilities.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Altclip: Altering the language encoder in clip for extended language capabilities

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.941293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.833993Z digest=sha256:1b0e2144e71bcd11f2809be9df1fbdac383aa98e67ea75a0c59bc046826b390d

Observation 818f101b-df01-47cb-9ef8-8d580dbea403 · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Maskclip: Masked self-distillation advances contrastive language-image pretraining

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.927385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.840322Z digest=sha256:d3427a5b2f1cbcbdb7ee876d2b128a08ab10a8e33b1efd56d660103196f6ae10

Observation d7465e84-03ff-4c58-ba74-e7446b4560ab · outbound

This paper cites An image is worth 16x16 words: Transform- ers for image recognition at scale.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training An image is worth 16x16 words: Transform- ers for image recognition at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.906995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.845096Z digest=sha256:4da9372e7c0981931553e2793aee3fa08b6a9bae155efba06b0f248be1a6680a

Observation 10bec2c1-eddb-469a-9391-df7cae57fefa · outbound

This paper cites Improving clip training with language rewrites.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Improving clip training with language rewrites

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.891952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.849417Z digest=sha256:afd79a436517c7cc6bd54ed38707c764e61fb53c01c87aa063cc8ae23915bdcf

Observation 2dc147cf-df78-40ce-8f07-695493844f6b · outbound

This paper cites Pyramidclip: Hierarchical feature alignment for vision-language model pretraining.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Pyramidclip: Hierarchical feature alignment for vision-language model pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.853452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.853452Z digest=sha256:d8f299a6873d816ed2e9e8e4c8f3d67dc7d377f467f26da1c8eaa042acd9fe59

Observation f562151e-0523-4909-994a-422254de4be5 · outbound

This paper cites Softclip: Softer cross-modal alignment makes clip stronger.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Softclip: Softer cross-modal alignment makes clip stronger

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.873648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.857225Z digest=sha256:1da2a67214c39e705bf39383b8bad4aaf6de2d16aaed23d9b6e243c793f18b64

Observation a44e35a2-0ca5-42d9-9767-00719897ee77 · outbound

This paper cites Hiclip: Contrastive language-image pre- training with hierarchy-aware attention.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Hiclip: Contrastive language-image pre- training with hierarchy-aware attention

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.861679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.861394Z digest=sha256:0bbcee975cf3979b4503dbcf901a94798d9acc27ed253e4e037e783843b3bebd

Observation f8b962ac-6dc7-4488-a66d-56777c37f4a0 · outbound

This paper cites Sugarcrepe: Fixing hack- able benchmarks for vision-language compositionality.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Sugarcrepe: Fixing hack- able benchmarks for vision-language compositionality

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.849945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.865235Z digest=sha256:ef42e59002a50149eff13eaa9c6ed5e406177e7b5a873475a5efc5fc63634053

Observation 8a850156-b50c-410f-a52f-a7e55f11e682 · outbound

This paper cites Openclip, 2021.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Openclip, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.837127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.869235Z digest=sha256:aeffc426ed340c390d4872e57db90117d9c681444efcef999ae90b120bc5196e

Observation cd64749d-6852-4338-bceb-c524d5dfda4c · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scaling up visual and vision-language representation learning with noisy text supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.824649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.873432Z digest=sha256:14df0a97ec45f2af409406ff5a05f4b609e24d7da294c97c00fb23847196bfc7

Observation 00cb0a3b-3899-409a-b0e7-2a1a893c0f8f · outbound

This paper cites Mistral 7B.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.878448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.878448Z digest=sha256:3732fcbe3ce978388453c8f0f5dbef11b8c3402cad2bced2bb25bad02554ace6

Observation 8d033562-e4e6-40c2-97d8-eb4b9b788d4f · outbound

This paper cites Scaling Sentence Embeddings with Large Language Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scaling Sentence Embeddings with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.882734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.882734Z digest=sha256:b0c057a81cffc108a61ff3bdeddcdf10ec5a3daa5340216614b0bdd7788457c7

Observation 9523231a-5fe3-433f-b323-a0f5e73eb07c · outbound

This paper cites Misalign, contrast then distill: Rethinking misalign- ments in language-image pre-training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Misalign, contrast then distill: Rethinking misalign- ments in language-image pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.812134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.886495Z digest=sha256:53dd7a0890345813fb746aa6c49f27e59bd0c3a4fa60be7753b8185b71df3b2a

Observation 03a6eaf3-1f33-47d8-b5fd-cbadd5602711 · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.890177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.890177Z digest=sha256:a2111431151a6760a453011b8e26cd4b55e6fbe3164c2c9d9d26af53c741d41b

Observation df8046fa-aa08-4e92-9781-8bcb8859ac51 · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.894281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.894281Z digest=sha256:e5ac9e22e20be08a3b11c6b93779ceb6f876758a06bd4ae9ded66eff47f0702c

Observation e468e0e3-9f0b-4f11-9434-debbc4ba3d07 · outbound

This paper cites Uni- clip: Unified framework for contrastive language-image pre- training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Uni- clip: Unified framework for contrastive language-image pre- training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.800689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.898476Z digest=sha256:aadaeea0a0105f6dfbac98c9d1036abfd50307cea4db7b967dd3894f87d15fe1

Observation e9c16874-4ff8-4c72-9d0b-036dbe818bdc · outbound

This paper cites Gecko: Versatile Text Embeddings Distilled from Large Language Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Gecko: Versatile Text Embeddings Distilled from Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.902333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.902333Z digest=sha256:f744e72999e824403cb27e1712cc18250b80c8152ca357b24a18a075c14f995f

Observation 10ee5516-cff1-45b8-be52-193017adfe4b · outbound

This paper cites Meta-task prompting elicits embedding from large language models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Meta-task prompting elicits embedding from large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.788710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.906723Z digest=sha256:9b3210a1102f7c557bc44eea8cdae6214c954d735a837f1e337330bf9dfec981

Observation 6eccc694-931a-4ab0-8990-fbe6fecc8037 · outbound

This paper cites Scene graph generation: A comprehensive survey.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scene graph generation: A comprehensive survey

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.775909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.910853Z digest=sha256:6f087486ddc3bc4b59239b904c539e4be94176a1d13844fb00b0a7dadad6cc8c

Observation 16d083a4-7b19-4f7e-861f-e3aa071eb484 · outbound

This paper cites Grounded language- image pre-training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Grounded language- image pre-training

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.764217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.914822Z digest=sha256:6cdd90172a6fa86be53441b2505fac0915a763792e4fd7c8c8d7ec22548443bb

Observation c291c074-5c93-4778-b7b3-73303cae886f · outbound

This paper cites An inverse scaling law for clip training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training An inverse scaling law for clip training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.752353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.919615Z digest=sha256:1e65964df8adc38d28448a60378d7dae2473dc5f8c81b113935f7e56cb85896a

Observation ad7acedf-ad08-4128-aa9b-576f699a806f · outbound

This paper cites Scene graph generation from objects, phrases and region captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scene graph generation from objects, phrases and region captions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.741521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.923463Z digest=sha256:70e0d2bb1ed627ec6df9db70d126db0024c40ca5c26158f05ab855362248a645

Observation a00fbd1f-04a6-4e30-bb3a-28cd27b1015b · outbound

This paper cites Supervi- sion exists everywhere: A data efficient contrastive language- image pre-training paradigm.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Supervi- sion exists everywhere: A data efficient contrastive language- image pre-training paradigm

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.729879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.927328Z digest=sha256:216cd5cae4c063ee78c808cd5ad88bb7aaf43e5bbb7d8db792135f46296a85f1

Observation 27244eb8-92a5-4a14-9c51-e2409d99e880 · outbound

This paper cites Scaling language-image pre-training via masking.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scaling language-image pre-training via masking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.931234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.931234Z digest=sha256:b86b8045dd8a7c23cfb91d1971ab2fc7967d01205afb4c8e61143954799c90e4

Observation 7358f3c8-a38c-4ecc-ac23-1f39083baa42 · outbound

This paper cites Language quantized autoencoders: Towards unsupervised text-image alignment.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Language quantized autoencoders: Towards unsupervised text-image alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.709761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.935183Z digest=sha256:2bc6654b11a8b1d4203f5c346460d5e3cc5cbaf7ba87bcd45504a32e5db3eeeb

Observation 07d857dd-9ef6-473b-aeff-8efc72d12d18 · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training MLLMs-Augmented Visual-Language Representation Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.940013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.940013Z digest=sha256:9b3559bcef7d5936ddcd17a29c7b6ebf1825809a3aa24ec3a8c6065c8b548b13

Observation eb291ae4-5fcb-422f-bd0d-46559506bd91 · outbound

This paper cites Decoupled Weight Decay Regularization.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Decoupled Weight Decay Regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.944402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.944402Z digest=sha256:fe0af0fe208a0bec7b3a983da899bb1d1cfcaa5c28d60161507579dc7aaa698a

Observation 0686173c-cdf0-4d4b-b1df-3615ec2ce85e · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Slip: Self-supervision meets language-image pre- training

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.696620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.948730Z digest=sha256:0a04d757085e86f960ef1acc68a534698ea7335b6ed4c36e0325f5243c2f1bb4

Observation ecf32336-bbed-4254-bbdc-e17cb97b2bcf · outbound

This paper cites Generative Representational Instruction Tuning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Generative Representational Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.952619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.952619Z digest=sha256:1f55bf079c600bf0dc62896aabc8960ac466707574810d645a4cb669bc828041

Observation e9fcf66c-1c73-49af-a6c7-910c486ad503 · outbound

This paper cites Docci: De- scriptions of connected and contrasting images.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Docci: De- scriptions of connected and contrasting images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.683887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.956758Z digest=sha256:24021eafd450d5aa43ee41a156e68bfbf09f34c3202ee20118fc1cc013eeb743

Observation 37cbdc5f-e82c-4926-8245-1520109e40ca · outbound

This paper cites Learning transferable visual models from natural language supervision.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Learning transferable visual models from natural language supervision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.670928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.960633Z digest=sha256:95870813e99a0882c21823c30b8e3f628be5d5e84fc05bfdb64b098bbee871ac

Observation 5d1b15dd-82f4-4705-b955-ae8943f16fcc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training High-resolution image synthesis with latent diffusion models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.964348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.964348Z digest=sha256:0ee9b91e3c2b5730e2c2023ed46a7009a49b32df85bae760542fc4a02c580616

Observation 00ea4c14-7c8a-47b0-b0eb-3be3273f8e55 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.648904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.968709Z digest=sha256:bd3ba21438f66ff1644126a31e75ef74fcff4191d22af97b7ac9dbd414f2b108

Observation b94bf483-407b-4f26-b040-3338f133aeb3 · outbound

This paper cites Repetition Improves Language Model Embeddings.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Repetition Improves Language Model Embeddings

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.973434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.973434Z digest=sha256:c75b67225731a73332f060c4e91e6399985f90e961e66ac4cccfc463f63dd661

Observation 1b1a5244-7d1c-4258-b856-f9a6ca575b3c · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.979973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.979973Z digest=sha256:a1747e079ee97ff4db3729ea49cf88335ee903fcbf697329368f67403e91202d

Observation 64e4cd87-453e-4e07-aac4-042fc769be0f · outbound

This paper cites Crossmodal-3600: A massively multilingual multi- modal evaluation dataset.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Crossmodal-3600: A massively multilingual multi- modal evaluation dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.635017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.985183Z digest=sha256:4ba2279e717aecb93891dce744758acf7f65fd953c9b721912b72f34229ee03f

Observation a603a36f-318c-49e6-9fb0-e34cf50d15a4 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.620165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.989037Z digest=sha256:5653c522c262a86d7ef67bf8e5e55d904343ea0d5ebba4a716c2497245abc71f

Observation 49222733-a491-46c9-accc-61f176f490a4 · outbound

This paper cites Stablerep: Synthetic images from text-to- image models make strong visual representation learners.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Stablerep: Synthetic images from text-to- image models make strong visual representation learners

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.608271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:18.992900Z digest=sha256:d120398f3e58772743ae9df74d1a8f79395f732edbc38a67d3acf401e1f62dbb

Observation 3d212d76-ad4e-4472-a9a6-86c71f9e7b44 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.997000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.997000Z digest=sha256:79a64b6edbc4f9d1d9b12f3d54af2c54d034a99a352a0e4ee94bc58c8e5a5dfb

Observation 181ea0ff-5cca-4aec-aaee-a07ffa20816e · outbound

This paper cites Multimodal few-shot learning with frozen language models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Multimodal few-shot learning with frozen language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.596000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.001300Z digest=sha256:ec0567d3e14bdf0c14996ad1f86da8bf230e155aa36cef493223d579f987415a

Observation ee888382-854d-4e60-8872-239e223ef6e8 · outbound

This paper cites NLLB-CLIP -- train performant multilingual image retrieval model on a budget.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training NLLB-CLIP -- train performant multilingual image retrieval model on a budget

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.006394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.006394Z digest=sha256:d86a391e4fdc55fb66489b14cf018a43fd8370d645817242b356d90d51bcd548

Observation b07fbd1e-0b69-4a54-bc5e-52af05dbb357 · outbound

This paper cites Improving Text Embeddings with Large Language Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Improving Text Embeddings with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.011691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.011691Z digest=sha256:11b9c19984da68b7a796704aa0705fd77e004c90eddff7d3c4bdaaf2a385d3b0

Observation ef407a48-25e3-4da7-8ac7-cbc02f5c9153 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Groupvit: Semantic segmentation emerges from text supervision

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.584093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.017503Z digest=sha256:099880ce7e566ee6f94432fd03fb674e11a150fb6fd2f70b7ea98b38d39aea19

Observation 5ba97f96-d30f-4d48-8a22-a1c68accd075 · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.021241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.021241Z digest=sha256:5dccaf00879f3a5edb3184afe28c2a0b9b9fed9e79256c46394eab3aa819f286

Observation 0d6bbb8b-f28e-4df6-b230-c09b5f4bffa4 · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic caption.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Alip: Adaptive language-image pre-training with synthetic caption

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.572338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.025421Z digest=sha256:2469a47a8da2d09afd883cca75663cf41237c977537e8482178b8b9496fa0db1

Observation 08a0c0fb-f71e-41cd-9394-4ac6a690eb78 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.029722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.029722Z digest=sha256:c7bfd2f983d58a9270ffdd65ba9371f01b4a380171304df839ea290e85270799

Observation b9688f0f-3c42-4e1c-8b41-f20410f578c1 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.034360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.034360Z digest=sha256:e291c647bd2699efd7b55c15d3d77868a194bf58a2b529cdc3be257d4ed7b380

Observation 6700abaa-d14e-4096-b8d4-80892ca4f79f · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.456970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.038843Z digest=sha256:049849cf81598f7d57cb6cb080cb5ff9a783751b12e7361994244f609ca53632

Observation 46fbb959-e44a-465c-a66f-4232af836c81 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Lit: Zero-shot transfer with locked-image text tuning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.445829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.042838Z digest=sha256:9a6cfd16cd18a78c894c2ca3cfcdfd5bcb52a9adb29614d80d02abb962bc7c3c

Observation 24684c16-7b27-4655-9deb-003400c7547c · outbound

This paper cites Sigmoid loss for language image pre-training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Sigmoid loss for language image pre-training

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.433717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.046846Z digest=sha256:d66f275aa37f0c79d20b208b8387ca1d1e6ff380c9576b36095f7752809a3205

Observation 0a8f4552-a5fc-4809-bdf8-9683dbfab0d1 · outbound

This paper cites Simple techniques for enhancing sentence embeddings in generative language models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Simple techniques for enhancing sentence embeddings in generative language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.420132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.051036Z digest=sha256:da9189f2a3e175a395e8a752de24feb2d46c56960c54fcd129b779248594b635

Observation 688377fb-d2d8-4516-b32b-70dc64e34dc9 · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Long-clip: Unlocking the long-text capability of clip

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.406652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.054747Z digest=sha256:1a3b1f97b99bb04d7aff78500c37391247d0a209564581d0d377e013663d6221

Observation d5678b42-79b4-4865-aee6-2f0a69950a5f · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Dreamlip: Language- image pre-training with long captions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.394875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.058927Z digest=sha256:453f2f849d00015bbcee4084516385eb971aadbe865c1b4290c9107530451787

Observation 31db733f-781e-4cc3-a199-1dcc8e19c9c8 · outbound

This paper cites Beyond text: Frozen large language models in visual signal comprehension.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Beyond text: Frozen large language models in visual signal comprehension

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.382821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.062778Z digest=sha256:8be473cd4e9b71331d7183e89a5fdf74888c3e2f65f0baf581fa96862f792c49

Observation bbb638db-a5df-4c67-a476-1f77e401f0d7 · outbound

This paper cites yi”. After thinking step by step, the category of the main object in this image means in just one word:.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training yi”. After thinking step by step, the category of the main object in this image means in just one word:

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T18:38:19.370466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T18:38:19.066615Z digest=sha256:2d677ab75754b9b4bedff003f9a778c17bcb656c952c7b8301ebe4db8f9c0854

Pith citing papers

Observation dccbc202-eb71-431c-a90a-e929ada7b54e · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.578691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.578691Z digest=sha256:3bdac829dcf760ed094c38e9ab251b661b0be205c6b7fd4fbc20b2ab08b841d7