Pith. sign in

Paper Citation Record · LEDGER

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2506.03096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03096 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:28.606467Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:55.798269Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:55:29.193720Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved15
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed39858-fade-4067-98e7-e06caba8e887 · outbound

This paper cites 4M-21: An any-to-any vision model for tens of tasks and modalities.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M-21: An any-to-any vision model for tens of tasks and modalities

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.557399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.087022Z digest=sha256:b2f04f437044e7451e2f1b77d8b19d1fb42ceb9e381a580516cc72d2d20a677b

Observation 9cb93a64-e857-4b93-9163-3b379ce22227 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Zero-shot composed image retrieval with textual inversion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.296341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.178855Z digest=sha256:9d3229023e00810715c2d4dcb24981b089f94c52ecadb4b94c650fa503cf4b6f

Observation 843f6fe7-90e7-4cd0-9951-977d031bbeda · outbound

This paper cites BEiT: BERT pre-training of image transformers.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens BEiT: BERT pre-training of image transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:34.129683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.280121Z digest=sha256:7ee3ffb1c66ebdcb39525bfa2eb5e64dd140a71391c58c07a4d4cae0f45d33a6

Observation 1b869257-70c1-4cce-97e8-a88289041117 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:23.390643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:23.390643Z digest=sha256:85e282cc4afa4fc8205deeb84c31198285b19c33afb15ea277cc3f71558b1882

Observation d894bb48-436a-48e4-ba01-d7f94534e70f · outbound

This paper cites All you may need for vqa are image captions.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens All you may need for vqa are image captions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.990034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.517383Z digest=sha256:465e6ee29aa3415e3cfec3be81745f06a23ebe1199190c22067baaa548db3124

Observation 494f61e7-b2b6-47b9-9471-a78c51515c42 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.886746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.580996Z digest=sha256:29139c2628948ed4f65d1d5f0afd98ab8588a6e7a127dec1dd0fc3411981c6c0

Observation 08e897f4-988d-438c-bc76-ef9b4c3b8f12 · outbound

This paper cites Understanding transferable representation learning and zero-shot transfer in CLIP.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Understanding transferable representation learning and zero-shot transfer in CLIP

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.775659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.639244Z digest=sha256:03d2ea3b629a1dacf06fae9b45182d03296b4f8277607676ee2f85672f8c61f6

Observation 1cf73c11-e53e-4c18-9520-d7be9186633c · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Reproducible scaling laws for contrastive language-image learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.682817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.759334Z digest=sha256:b5c420ba0e91dc8ed0f93cfed28b02351b9ef522a88ed21198892766370e31b7

Observation aa8602d3-d47b-443a-b5d2-dcfca48a8565 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.557702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.837170Z digest=sha256:ce8a48129b2a7650a57dac02f5cd9dcb47c337b90e031a5dc6152600c8c12832

Observation cb770c8c-2a79-4e1d-adb2-3049e3c59a6a · outbound

This paper cites Bert: Pre-training of deep bidirec- tional transformers for language understanding.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bert: Pre-training of deep bidirec- tional transformers for language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.464984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:23.956817Z digest=sha256:bd8d8403cf0456f0a2a53dbb0214c6d292003713761bbac8dc9937a6eec35675

Observation c432fe62-658c-4d53-ad9a-a9022fc97142 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.354101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:24.060445Z digest=sha256:7b5d15b8daaf8765798aa1974935ec6e15e686550aff22c64145ee4683d8ec3b

Observation 0fd1af9d-39f7-4783-9788-f1d0f4d9fbdb · outbound

This paper cites The Llama 3 Herd of Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.193921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.193921Z digest=sha256:69bf6bdd8a49a20b4f1bb9f7198c0d2f0ab149522c2091a869585a4859b29999

Observation b11d3bb8-0ab1-4fa2-8f9b-27fc5fc7f308 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Taming transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.245457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:24.311637Z digest=sha256:39d6267e760d2078dc2bbe6f845bc5b32f7e5b5103d56ee575f1667b35034d8a

Observation bb475763-a59d-427a-b4c8-7ae4b4a5f9c4 · outbound

This paper cites Dreamsim: Learning new dimensions of human visual similarity using synthetic data.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Dreamsim: Learning new dimensions of human visual similarity using synthetic data

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.129797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:24.448475Z digest=sha256:ce9d2cbb76c5692ee6e3dc250b86a08dcdf6b3904bbedaacdcd85bbce2be252e

Observation 7e2d01de-0b28-4cb6-8211-104430daa9f5 · outbound

This paper cites Language-only efficient training of zero-shot composed image retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Language-only efficient training of zero-shot composed image retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:33.021382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:24.584746Z digest=sha256:55c2e03ef69dde19a8b41778c7a58b8193933fbf8bb4e171cd266bc8c098c423

Observation 837e14ff-b602-40d5-b2fc-b2d9550d026a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.919515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:24.674048Z digest=sha256:c6b93fba728a4d001a10857d3ce182cf39f2e3248efdc7eb3e644058ff6726a6

Observation f5f16587-952d-4cf0-b754-d2c365832953 · outbound

This paper cites Hq-edit: A high-quality dataset for instruction-based image editing.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hq-edit: A high-quality dataset for instruction-based image editing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.837374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:24.789096Z digest=sha256:445b1a04e1b33bbfc8fba09d510fc8b49ddf647ef17860f44f987371964930c5

Observation df0adcd8-bfce-482c-a84c-7dae883e1ac1 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Scaling up visual and vision-language representation learning with noisy text supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.894088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.894088Z digest=sha256:0d0584fa2fea495464aa658bbceb0c6b3ac4e47f9b9e1eeb6914211aceed8f94

Observation e334eac3-8f55-4ee9-8da4-99fb353502f9 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:24.985554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:24.985554Z digest=sha256:042985e270daacdaaba1e04d97c19a9988039594e9ff54ab615b72b63cf294b9

Observation b801fa92-b8c0-4041-8bb0-87e2ce385ef8 · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.716035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:25.098746Z digest=sha256:65d371680f04f65f72a39205c1c24fadfe92561d818da44a8e56420f35a3d7d9

Observation 5c55ce28-7e71-442b-9775-b612e8373450 · outbound

This paper cites Hard negative mixing for contrastive learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hard negative mixing for contrastive learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.613227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:25.212454Z digest=sha256:a6904f9f38512e5b0153bb943374c955d208c21d46013d1287c51c548cef3c21

Observation b7f3d2fa-61a0-499e-b276-ddfe51cae974 · outbound

This paper cites Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:25.297748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:25.297748Z digest=sha256:c591ffb514e24ac484fe46435352417e6d62ec7326a364e14363e8c013585534

Observation 163d1407-a624-4fcd-ba8f-b3bfae928cfc · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual genome: Connecting language and vision using crowdsourced dense image annotations.IJCV, 2017

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.525438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:25.415308Z digest=sha256:003d2cd99cb040eeab03f58e218e7617facb49aaf7fac8e15d3e36146942e3fb

Observation 3d965c66-e85b-4216-a5ae-0e563dfeb644 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.417980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:25.518996Z digest=sha256:7392ddebdeb4e3b5b670e728fb518d1a21fcb33d576ff95e77e4d08b95f4e83a

Observation 7d89c8a0-66b4-4823-8681-f9057a169c61 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.304284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:25.630094Z digest=sha256:6ea35d67c5762830fce5f086028bd14078f757bc198390aad77df52a98b45cd0

Observation e07f0e3f-31d0-427a-8f86-cc262e0424ec · outbound

This paper cites Visual instruction tuning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:25.764193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:25.764193Z digest=sha256:1539921306d91c6bf0a988115ffeaa62a83dd0659249243f0923e20737806bad

Observation 9267b3c3-057a-4a59-b17b-dc84985ae7ee · outbound

This paper cites Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Universal vision-language dense retrieval: Learning a unified representation space for multi-modal retrieval

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:32.155301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:25.869859Z digest=sha256:46d6b712de2d3092b483fe77079e9f4e845d146a3371e7cc56d9252dfed1fe69

Observation f831d799-616a-4c29-a57a-844c122c90c8 · outbound

This paper cites Decoupled weight decay regularization.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Decoupled weight decay regularization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.992247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:25.953221Z digest=sha256:3407a37d3a312d61cb542c9e4e38d4da9882867a28910e21068c3b87117ca91a

Observation 0330e435-b54c-45b0-8640-3fb3636ff3be · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.091512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.091512Z digest=sha256:b50372f9f3a9847aaa0b9ebccedfb259ec566ab77b94778518b5e32b70adb35f

Observation ebb90a6d-3e34-4faf-94fa-06e656d78717 · outbound

This paper cites 4M: Massively multimodal masked modeling.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 4M: Massively multimodal masked modeling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.867994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.267694Z digest=sha256:9181001ddd9d398558781bdb0ef9af7d3aa3dbddb348cd2ef8a4c7157517aab7

Observation 5ca6bf1b-f61f-42b0-9b94-3fb20cabc4f3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Learning transferable visual models from natural language supervision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.778503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.367835Z digest=sha256:a20785af8b1c346ffef2b7f8c3d3623c104fd501067cad6292f1238d4d2b6c09

Observation 685d25df-f3bd-4182-9146-cbfbb3299bdf · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.649413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.479501Z digest=sha256:99b7a5ab8ec5b5c30d6bfd38da503f8d20888bd88f5b481e170f9df3c8c38282

Observation 726aa4b1-bdb6-4c82-bcfc-c29f883b32f8 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.584738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.584738Z digest=sha256:d9861557bb7af65f2abd1bb3cde43c2e4cfe61058d5c3fa6a1c57445d8c0afb6

Observation fe737f5e-fc12-40b2-bf4a-a691b7b3fc69 · outbound

This paper cites Contrastive learning with hard negative samples.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Contrastive learning with hard negative samples

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.556519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.665589Z digest=sha256:9a8cd684bdfe0edf805a8696db472f06dcf911b8607f7d609db75b031d955058

Observation a161c2c6-7052-44f3-b35f-02ce37ba2616 · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.418550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.746652Z digest=sha256:a8dff078b850c16fc12af8a2e5030dd6d100a0d674e172b2f4af860502e539dc

Observation 0b3b8330-ed67-48e4-be40-aaa806d0b0c1 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.299067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.802857Z digest=sha256:0e20a1af9f8c1dc04650577c26cf9feba171bf98de5ac03241ff758e3aeb5ae9

Observation e467b26d-a173-4250-b527-28093c29a0e8 · outbound

This paper cites Towards understanding the modality gap in clip.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Towards understanding the modality gap in clip

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.178468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.901538Z digest=sha256:efe239521baaf47a9b0e992b683347930f0f8f14a273c2c01f0592e9fc1f7f05

Observation 90f3f5d9-5e2a-4835-8190-6877d5b1b812 · outbound

This paper cites FLA V A: A foundational language and vision alignment model.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens FLA V A: A foundational language and vision alignment model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:31.046467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:26.965593Z digest=sha256:480d05ad6ccb5568474d7e2fe607fd980ecf11951f495aca6d10ba0f2468b53b

Observation a10a96e6-01a0-4da0-9c8d-1df34c23561d · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.042563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.042563Z digest=sha256:8d9375dbca54b167f39bfc057a7f0ab1c8fc7905f467e955f3905de25e55216d

Observation 613cdf8f-2a49-4ec4-8fa3-f44c17f4e834 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.136459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.136459Z digest=sha256:1f55072630b01cfeb2a500b512b641611dd597e7067e6378236bd809731ada3c

Observation 497cf7b5-a8b4-43e8-8fea-914f796a061b · outbound

This paper cites Neural discrete representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Neural discrete representation learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:27.251290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:27.251290Z digest=sha256:4ab85c2db38e955f1c9c8524f9958a9ffd3388bad5ed95b9541e3f97111e277d

Observation 05ee0f4a-8ca8-4684-bd1b-1a62196deba7 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to- sequence learning framework

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.869888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:27.350816Z digest=sha256:3e8ae625df73e625fa9186b2178b08bef33d06726532c942d484ee021397e124

Observation 505e623d-3c28-4808-bf86-62e062430458 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retrievers.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Uniir: Training and benchmarking universal multimodal information retrievers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.710665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:27.482449Z digest=sha256:c1bcc09882f9896ef68f6c7a6e9ae933a72304211c76eab57638608d2f51d7ce

Observation 112ca790-d670-4692-8f41-5d4b31f18c51 · outbound

This paper cites Robust fine-tuning of zero-shot models.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Robust fine-tuning of zero-shot models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.570571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:27.604662Z digest=sha256:5456b04b2ef28a4f6f7fc25e922413a2104074f10e72a5fba78bf9f9b5228d7b

Observation 9d2c0216-9f12-4504-8b26-37d8d0e3c5f9 · outbound

This paper cites Bridgetower: Building bridges between encoders in vision-language representation learning.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Bridgetower: Building bridges between encoders in vision-language representation learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.449384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:27.701150Z digest=sha256:e7c6cb1bbb292a90b26972f6bf9fd59c1bf04d686a6c7ef5442cbb6fe9d0b1b8

Observation f278eb64-6d5f-4f50-ab75-56c4c614c8fd · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens An image is worth 32 tokens for reconstruction and generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.310541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:27.794105Z digest=sha256:502d44b761a929d607e12853c9119739ea184192c9b9292b2c7fa1b75aa85e97

Observation 01dc3f6d-aaf4-4c64-b5ba-ac51523ae568 · outbound

This paper cites Sigmoid loss for language image pre-training.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Sigmoid loss for language image pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.160191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:27.896708Z digest=sha256:6523faf75cc3e77c7dd72e7701417ef87483347a42329ffa241ed26511eae9a6

Observation d637248f-3224-45c9-ab61-2d67cb6366af · outbound

This paper cites Magiclens: Self-supervised image retrieval with open-ended instructions.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Magiclens: Self-supervised image retrieval with open-ended instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:30.023795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:27.988902Z digest=sha256:ce51451a8440fbb40734d21cf81c05405df3eb65ea9b5cec9c247e75e872c336

Observation d2b67434-2652-493e-b0bb-6da00b6f1b67 · outbound

This paper cites upper left, upper center,.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens upper left, upper center,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:14:29.864818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:28.101210Z digest=sha256:23cffacf67382a1a0ad592fc9daa03e5665ca21ebbdf9206266f6adda87e140f

Observation 87177a54-1e14-4952-9ab1-f70f33b5a5aa · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.704991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:28.138032Z digest=sha256:beca237d12740811417b5e42c8e333f7ebd457996a13035782e17f885844430c

Observation 484e56ad-241c-484d-8993-449bba093f24 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.574570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:28.217502Z digest=sha256:558c9f27d56737f5bab27bf135981398f0a44caa3fc92e7d018a51901ffbf459

Observation 7691d229-b6cf-4efc-8da9-85f6e5f3b099 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.421030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:28.321253Z digest=sha256:d48fa6d6a4040218150ab0fcd4780fcfb66957d06ef75996b84cd806ddb96a2a

Observation dc538a3b-e078-43bc-9b44-b08d22f7d502 · outbound

This paper cites an unresolved cited work.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:14:29.252149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:28.404311Z digest=sha256:bc2007caa5e307468e02d78ead1e77d88572b3e3f024d128fe728546f5d835c5

Observation dd478e66-6fbf-4a1c-81b5-6c79b2bb817e · outbound

This paper cites The {object_name} on the left/right.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens The {object_name} on the left/right

Reference 54

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:14:29.054862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:28.511265Z digest=sha256:33c20385c730e2e381f2bc92984f472c47f7ebea4dd0e1184236acb5530e2f42

Observation 53a2cc21-01d7-4499-973b-e6ec13121e93 · outbound

This paper cites Notably, this model is much larger in the amount of parameters (4.15B, i.e.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Notably, this model is much larger in the amount of parameters (4.15B, i.e

Reference 55

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:14:28.889029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:14:28.606467Z digest=sha256:b5a73abb5023029b72998d9ea0ef68f75211b11d52302d4f3904bbe8fe974313

Pith citing papers

Observation cdce7db5-aaeb-44ed-b0dc-df797bdead84 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:55:29.196787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:55:28.067052Z digest=sha256:1882abaee65da7246ebe328b7ca0e0fa276cbb33b7202ebac3b8a91b25fdcc42

Observation 3af209f2-cc82-4d58-880e-edb5f2e87fb5 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:55.798269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:55.798269Z digest=sha256:0e305354d1f763ff84c05c94c5f6e7b103ffef51efa6903bd9c58d5d80d28724