FuseLIP processes image and text tokens in one shared transformer, trained with contrastive and masked-modeling losses, and shows that early fusion can beat late fusion for multimodal embeddings.
Zero-shot composed image retrieval with textual inversion
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
FuseLIP processes image and text tokens in one shared transformer, trained with contrastive and masked-modeling losses, and shows that early fusion can beat late fusion for multimodal embeddings.