Pith. sign in

Paper Citation Record · LEDGER

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2506.03096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03096 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:28.606467Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:55.798269Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:55:29.193720Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved15
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed39858-fade-4067-98e7-e06caba8e887 · outbound

This paper cites 4M-21: An any-to-any vision model for tens of tasks and modalities.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M-21: An any-to-any vision model for tens of tasks and modalities

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.557399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.087022Z digest=sha256:83db8e07ec516ab7bdec9d1bca0fed9100cc211d9783b115c7bcd8be3af2b812

Observation 9cb93a64-e857-4b93-9163-3b379ce22227 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Zero-shot composed image retrieval with textual inversion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.296341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.178855Z digest=sha256:52feb8ffa49e10fb179e85c28b9b212b71bb279bb7f590949e6de65972c0d586

Observation 843f6fe7-90e7-4cd0-9951-977d031bbeda · outbound

This paper cites BEiT: BERT pre-training of image transformers.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens BEiT: BERT pre-training of image transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.129683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.280121Z digest=sha256:8c446fe21e374705fb56520bc45a6348e1024ef9fe06264e50e24857aed5e8e0

Observation 1b869257-70c1-4cce-97e8-a88289041117 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:23.390643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:23.390643Z digest=sha256:eb7f614410e488ccb3ea22d892316ab3bb2a984fd288d4341af3fd9ed984a485

Observation d894bb48-436a-48e4-ba01-d7f94534e70f · outbound

This paper cites All you may need for vqa are image captions.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens All you may need for vqa are image captions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.990034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.517383Z digest=sha256:b96c03dee589fb3d9bd53b2f168eb32e3c68d6626b45bc75d259c1d7ff083084

Observation 494f61e7-b2b6-47b9-9471-a78c51515c42 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.886746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.580996Z digest=sha256:6c70ab3ebfc4f53fb27c817485cae32c6f8f3cc25523207f0d10ae39c3069ec8

Observation 08e897f4-988d-438c-bc76-ef9b4c3b8f12 · outbound

This paper cites Understanding transferable representation learning and zero-shot transfer in CLIP.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Understanding transferable representation learning and zero-shot transfer in CLIP

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.775659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.639244Z digest=sha256:14495493b81b0e5684df91741247c3956b19ac71105f27fd4927d20322d854a0

Observation 1cf73c11-e53e-4c18-9520-d7be9186633c · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Reproducible scaling laws for contrastive language-image learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.682817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.759334Z digest=sha256:180991b8423181e0723d40b5686c0e4ed14947bef3ca0af16d62b8a53dcc1e9c

Observation aa8602d3-d47b-443a-b5d2-dcfca48a8565 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.557702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.837170Z digest=sha256:81e49413404a56b9430c3cd85da9cc80ff9ffcba3df183bf67997854685e5c33

Observation cb770c8c-2a79-4e1d-adb2-3049e3c59a6a · outbound

This paper cites Bert: Pre-training of deep bidirec- tional transformers for language understanding.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bert: Pre-training of deep bidirec- tional transformers for language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.464984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:23.956817Z digest=sha256:4fdfe7cbca8989cc5f770b82701448213477cae2ca6235a3846bb002f57440a9

Observation c432fe62-658c-4d53-ad9a-a9022fc97142 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.354101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:24.060445Z digest=sha256:c5566a0f0682209c3fbd19a6f27b6f200e66cec829ee82bfc5914bebe906be73

Observation 0fd1af9d-39f7-4783-9788-f1d0f4d9fbdb · outbound

This paper cites The Llama 3 Herd of Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.193921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.193921Z digest=sha256:7dbd0d82a6b5745688464cc948e71d01cf0c112e139dac44b9caa7fb1ae2d10b

Observation b11d3bb8-0ab1-4fa2-8f9b-27fc5fc7f308 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Taming transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.245457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:24.311637Z digest=sha256:dad30eaa51944c4e4b1fde0f205c0e8017a6c5cc71bc6305f7cf0c20bddcc4c8

Observation bb475763-a59d-427a-b4c8-7ae4b4a5f9c4 · outbound

This paper cites Dreamsim: Learning new dimensions of human visual similarity using synthetic data.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Dreamsim: Learning new dimensions of human visual similarity using synthetic data

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.129797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:24.448475Z digest=sha256:e5b6785b1e5bc8f37b475ddd2d7dab0897f737c506b31e64b8349cd2a6954368

Observation 7e2d01de-0b28-4cb6-8211-104430daa9f5 · outbound

This paper cites Language-only efficient training of zero-shot composed image retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Language-only efficient training of zero-shot composed image retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.021382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:24.584746Z digest=sha256:9fd67d8963d108950d2aa950684fe80f76c2093932a3ec150f54be5eaae2b9d8

Observation 837e14ff-b602-40d5-b2fc-b2d9550d026a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.919515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:24.674048Z digest=sha256:4fa3115c0a07902875388db0c6902538b946ded1c6de439dfbb6d83351033e82

Observation f5f16587-952d-4cf0-b754-d2c365832953 · outbound

This paper cites Hq-edit: A high-quality dataset for instruction-based image editing.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hq-edit: A high-quality dataset for instruction-based image editing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.837374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:24.789096Z digest=sha256:dcbfc496e61d4f6fe8c941a222aef2306d7a457d5b8bed9804de47878bc0ac55

Observation df0adcd8-bfce-482c-a84c-7dae883e1ac1 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Scaling up visual and vision-language representation learning with noisy text supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.894088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.894088Z digest=sha256:b12f485a48716e376c8dd206a96996fa3c219dd02d1354f5b269b1282df5a8bd

Observation e334eac3-8f55-4ee9-8da4-99fb353502f9 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.985554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.985554Z digest=sha256:c46cb2b876d72b75085cc0b9a56770f902872709fd0ef98b3e24741eb1679a7b

Observation b801fa92-b8c0-4041-8bb0-87e2ce385ef8 · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.716035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:25.098746Z digest=sha256:0346b0d019c22528a86aa653a0b5595bccd7625deb219cff8e3460e41289cf19

Observation 5c55ce28-7e71-442b-9775-b612e8373450 · outbound

This paper cites Hard negative mixing for contrastive learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hard negative mixing for contrastive learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.613227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:25.212454Z digest=sha256:9230771c2926a802094554b8348f3138c8edca1ddd3146a68d104de74fa8a5fd

Observation b7f3d2fa-61a0-499e-b276-ddfe51cae974 · outbound

This paper cites Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:25.297748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:25.297748Z digest=sha256:413f6794753d5feafa3c7f2c0c045898b7c4df0c7ed0dc5f040ad2ee79642abd

Observation 163d1407-a624-4fcd-ba8f-b3bfae928cfc · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.525438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:25.415308Z digest=sha256:ab5dc7878164ae5b95a943f2c1e08fcbe6fb113a96e276cf09d3ac74e1a2a931

Observation 3d965c66-e85b-4216-a5ae-0e563dfeb644 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.417980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:25.518996Z digest=sha256:cbf07dfd1308b539d2ad4d8e434276f722e22367590cc3625d34f98baf31bf4f

Observation 7d89c8a0-66b4-4823-8681-f9057a169c61 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.304284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:25.630094Z digest=sha256:61db8077bc6b353917149ac1a227354b1b09838057ce0ee072812f7a0bf14036

Observation e07f0e3f-31d0-427a-8f86-cc262e0424ec · outbound

This paper cites Visual instruction tuning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:25.764193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:25.764193Z digest=sha256:2dfc65d2cf50bbeb9a2a1f7fe57f2d7bc1db1a0fd173177810dd7c3cc0bf472f

Observation 9267b3c3-057a-4a59-b17b-dc84985ae7ee · outbound

This paper cites Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.155301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:25.869859Z digest=sha256:2e6d65bc4b5bba822bc4b63a4806c3433972449e16e139b00462d25e91af63df

Observation f831d799-616a-4c29-a57a-844c122c90c8 · outbound

This paper cites Decoupled weight decay regularization.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Decoupled weight decay regularization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.992247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:25.953221Z digest=sha256:23f74b4eaeb6133edfebd79829b413d1e080c06ac6e3b51f51b96b0d4a45a7e4

Observation 0330e435-b54c-45b0-8640-3fb3636ff3be · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.091512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.091512Z digest=sha256:d01e6d3ee1e0dd27230a14ccbece06ad9467b13830ef4683191cf6684e412faa

Observation ebb90a6d-3e34-4faf-94fa-06e656d78717 · outbound

This paper cites 4M: Massively multimodal masked modeling.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M: Massively multimodal masked modeling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.867994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.267694Z digest=sha256:8c8407bf025d68a4a2dad465e0816c9af8c117fb20b1d2e5146b6745544f4fa2

Observation 5ca6bf1b-f61f-42b0-9b94-3fb20cabc4f3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Learning transferable visual models from natural language supervision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.778503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.367835Z digest=sha256:56c0b50992b7394896420cbc4a10f20770897118c47c716a8bb262a0d048318d

Observation 685d25df-f3bd-4182-9146-cbfbb3299bdf · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.649413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.479501Z digest=sha256:ea7124135bf3ca238a2bffa32b7036608b607a808f9405971641ae36416f7b69

Observation 726aa4b1-bdb6-4c82-bcfc-c29f883b32f8 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.584738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.584738Z digest=sha256:75a261fbec4af259d25141aa16bbe4e61540112e8f193abcdfffd840aa8f7a08

Observation fe737f5e-fc12-40b2-bf4a-a691b7b3fc69 · outbound

This paper cites Contrastive learning with hard negative samples.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Contrastive learning with hard negative samples

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.556519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.665589Z digest=sha256:94a680c623730ae16af8bc86e9f43b06459856cffaf919a87e1606064485d809

Observation a161c2c6-7052-44f3-b35f-02ce37ba2616 · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.418550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.746652Z digest=sha256:6bc51afeab8b5ae3186ff1ddb9e58186337551adf344184ed8666027a1c0fe51

Observation 0b3b8330-ed67-48e4-be40-aaa806d0b0c1 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.299067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.802857Z digest=sha256:4a7ac38232fb7f67fb29e1b625a838d1f46759af56ff53289830ff9a95703f7d

Observation e467b26d-a173-4250-b527-28093c29a0e8 · outbound

This paper cites Towards understanding the modality gap in clip.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Towards understanding the modality gap in clip

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.178468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.901538Z digest=sha256:5b5bc4f1dfcdc7e08d5a73c683cfe7761de427a16520159a4e3f962b667f7403

Observation 90f3f5d9-5e2a-4835-8190-6877d5b1b812 · outbound

This paper cites FLA V A: A foundational language and vision alignment model.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens FLA V A: A foundational language and vision alignment model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.046467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:26.965593Z digest=sha256:82a99211122ef95322ffcccb865ccdf18d99b211183d7ddf84b22bb23b101b18

Observation a10a96e6-01a0-4da0-9c8d-1df34c23561d · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.042563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.042563Z digest=sha256:6decac2e86fa66612b9034a02abb78a7dda54b1fb78d70c207f9b160b380437f

Observation 613cdf8f-2a49-4ec4-8fa3-f44c17f4e834 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.136459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.136459Z digest=sha256:8d7e884ae01e0e71dbaa94135e4f23ad9f03df6e00df37c432ebab8fd3b8062c

Observation 497cf7b5-a8b4-43e8-8fea-914f796a061b · outbound

This paper cites Neural discrete representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Neural discrete representation learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.251290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.251290Z digest=sha256:867b6bfeb90f439f39807da3eea8b26b4dba553970c6904d639545643934ac9b

Observation 05ee0f4a-8ca8-4684-bd1b-1a62196deba7 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.869888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.350816Z digest=sha256:580a42115b4dccb67908654ca18651e80e9b9471e962128cab2ea21daad1c39a

Observation 505e623d-3c28-4808-bf86-62e062430458 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Uniir: Training and benchmarking universal multimodal information retrievers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.710665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.482449Z digest=sha256:c7d88f9e2a229c9e95ecff7c2ef86082128935f1c3e3f69e09a85835318a7820

Observation 112ca790-d670-4692-8f41-5d4b31f18c51 · outbound

This paper cites Robust fine-tuning of zero-shot models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Robust fine-tuning of zero-shot models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.570571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.604662Z digest=sha256:367cf5edd41fb99308a1c813ad94f014cc6b480d27f4daeda3008dd3ec324d1c

Observation 9d2c0216-9f12-4504-8b26-37d8d0e3c5f9 · outbound

This paper cites Bridgetower: Building bridges between encoders in vision-language representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bridgetower: Building bridges between encoders in vision-language representation learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.449384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.701150Z digest=sha256:f8a0f38fb1bedd174d34923350977bf6932dd6a8bb3fabad855f3a20ebf371af

Observation f278eb64-6d5f-4f50-ab75-56c4c614c8fd · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 32 tokens for reconstruction and generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.310541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.794105Z digest=sha256:042bfbe5da522c59aa24d89e23892ae27da04e62f932f7f59bf58e19e44ee451

Observation 01dc3f6d-aaf4-4c64-b5ba-ac51523ae568 · outbound

This paper cites Sigmoid loss for language image pre-training.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sigmoid loss for language image pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.160191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.896708Z digest=sha256:0d98a42f2dfe03dcf47aa84be9f58391b5a0e02391a59871d218708117ccf1dc

Observation d637248f-3224-45c9-ab61-2d67cb6366af · outbound

This paper cites Magiclens: Self-supervised image retrieval with open-ended instructions.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Magiclens: Self-supervised image retrieval with open-ended instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.023795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:27.988902Z digest=sha256:51b33e3ad7f2aea925704a50ededf648b5c8666ca1bf92dbf404b517a5c9bfa4

Observation d2b67434-2652-493e-b0bb-6da00b6f1b67 · outbound

This paper cites upper left, upper center,.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens upper left, upper center,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:29.864818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.101210Z digest=sha256:9c3b437eef501ea9daac273b7a2c9c27008cf873954f7d78dce71c3ed5f2d88c

Observation 87177a54-1e14-4952-9ab1-f70f33b5a5aa · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.704991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.138032Z digest=sha256:3ee1db6b94f725b9a7a6e7aa3aee4acb6e684abaa05035804bdd22c9f13df973

Observation 484e56ad-241c-484d-8993-449bba093f24 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.574570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.217502Z digest=sha256:d8f83d851fbc57ee6694c148f298737f3612a937261ba60d23261621616e6b89

Observation 7691d229-b6cf-4efc-8da9-85f6e5f3b099 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.421030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.321253Z digest=sha256:d4ee49308c1f930e0a2d8f111966c129e0a1879743593c5cea971aa304e8fbb1

Observation dc538a3b-e078-43bc-9b44-b08d22f7d502 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.252149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.404311Z digest=sha256:3541e49c10fc0d1bc825d2cc0a78cfd60ba716ede0c87e5fe4bd9600ed948f45

Observation dd478e66-6fbf-4a1c-81b5-6c79b2bb817e · outbound

This paper cites The {object_name} on the left/right.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The {object_name} on the left/right

Reference 54

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:14:29.054862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.511265Z digest=sha256:5906b248a4ecbf11e86f2c48dce154797ef64d849667ccd91d8366c9a73db62c

Observation 53a2cc21-01d7-4499-973b-e6ec13121e93 · outbound

This paper cites Notably, this model is much larger in the amount of parameters (4.15B, i.e.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Notably, this model is much larger in the amount of parameters (4.15B, i.e

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:14:28.889029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:14:28.606467Z digest=sha256:ca6756c94a35e36c6b51d1589582eeccc5bbb7a5d69b8e2898d3e40d9a8e89aa

Pith citing papers

Observation cdce7db5-aaeb-44ed-b0dc-df797bdead84 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:55:29.196787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:55:28.067052Z digest=sha256:748d19182ec4c9a2a00cb79ea053f38ef9a2562ac25f8ec350f15869ff4c91ed

Observation 3af209f2-cc82-4d58-880e-edb5f2e87fb5 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:55.798269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:55.798269Z digest=sha256:cd8c31137891fbcddc3a1dd1e7cc91d1c6b58f963bdbc9fbfe930acf701bf5e3