Pith. sign in

REVIEW 1 cited by

MATE: Meet At The Embedding -- Connecting Images with Long Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.09541 v1 pith:IV2T2ZXT submitted 2024-06-26 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords texttextsimageslongmatecaptionsembeddingsencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While advancements in Vision Language Models (VLMs) have significantly improved the alignment of visual and textual data, these models primarily focus on aligning images with short descriptive captions. This focus limits their ability to handle complex text interactions, particularly with longer texts such as lengthy captions or documents, which have not been extensively explored yet. In this paper, we introduce Meet At The Embedding (MATE), a novel approach that combines the capabilities of VLMs with Large Language Models (LLMs) to overcome this challenge without the need for additional image-long text pairs. Specifically, we replace the text encoder of the VLM with a pretrained LLM-based encoder that excels in understanding long texts. To bridge the gap between VLM and LLM, MATE incorporates a projection module that is trained in a multi-stage manner. It starts by aligning the embeddings from the VLM text encoder with those from the LLM using extensive text pairs. This module is then employed to seamlessly align image embeddings closely with LLM embeddings. We propose two new cross-modal retrieval benchmarks to assess the task of connecting images with long texts (lengthy captions / documents). Extensive experimental results demonstrate that MATE effectively connects images with long texts, uncovering diverse semantic relationships.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

    cs.IR 2025-06 conditional novelty 6.0 of 10

    Vela adapts an audio MLLM into a universal text-audio embedding model using 'in one word' prompts, in-context examples, and text-only contrastive training, outperforming CLAP-style models on retrieval benchmarks.

Pith tools