REVIEW 13 cited by
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
State-of-the-art retrieval models typically address a straightforward search scenario, in which retrieval tasks are fixed (e.g., finding a passage to answer a specific question) and only a single modality is supported for both queries and retrieved results. This paper introduces techniques for advancing information retrieval with multimodal large language models (MLLMs), enabling a broader search scenario, termed universal multimodal retrieval, where multiple modalities and diverse retrieval tasks are accommodated. To this end, we first study fine-tuning an MLLM as a bi-encoder retriever on 10 datasets with 16 retrieval tasks. Our empirical results show that the fine-tuned MLLM retriever is capable of understanding challenging queries, composed of both text and image, but it underperforms compared to a smaller CLIP retriever in cross-modal retrieval tasks due to the modality bias exhibited by MLLMs. To address the issue, we propose modality-aware hard negative mining to mitigate the modality bias exhibited by MLLM retrievers. Second, we propose continuously fine-tuning the universal multimodal retriever to enhance its text retrieval capability while preserving multimodal retrieval capability. As a result, our model, MM-Embed, achieves state-of-the-art performance on the multimodal retrieval benchmark M-BEIR, which spans multiple domains and tasks, while also surpassing the state-of-the-art text retrieval model, NV-Embed-v1, on the MTEB retrieval benchmark. We also explore prompting the off-the-shelf MLLMs as zero-shot rerankers to refine the ranking of the candidates from the multimodal retriever. We find that, through prompt-and-reranking, MLLMs can further improve multimodal retrieval when the user queries (e.g., text-image composed queries) are more complex and challenging to understand. These findings also pave the way for advancing universal multimodal retrieval in the future.
Forward citations
Cited by 13 Pith papers
-
MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models
On 190 tasks and a 12-direction cross-modal diagnostic, seven embedding models frequently fail to honor explicit target-modality instructions: retrieval is biased toward the query modality and instruction-induced shif...
-
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
Using a structured rewrite instead of CoT as the generative interface improves MLLM-based multimodal embedding performance while cutting thinking tokens by about half.
-
ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding
Recurrently reusing the middle retrieval-forming layers of a multimodal encoder, plus small learnable registers as persistent state, improves universal embedding retrieval at much lower latency than autoregressive rea...
-
FreeRet: MLLMs as Training-Free Retrievers
FreeRet enables pretrained MLLMs to act as training-free retrievers via semantically grounded embeddings and reasoning-based reranking, outperforming models trained on millions of pairs on MMEB benchmarks.
-
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
A two-stage training recipe that converts causal VLMs into bidirectional multimodal embedding models, achieving SOTA on MMEB.
-
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
A systematic empirical study finds that for multimodal RAG, EVA-CLIP retrieval, listwise LVLM reranking, and feeding only the top-ranked document works best, with a self-reflection agent adding further gains.
-
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
UNITE combines curated multimodal training data and a modality-masked contrastive loss to achieve strong retrieval performance across text, image, and video tasks.
-
UniCoRN: Unified Commented Retrieval Network with LMMs
A frozen multimodal LLM is extended with a retrieval adapter and an entity adapter to retrieve a relevant image and generate a supportive textual comment.
-
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
Visual token compression (4x fewer tokens) plus a three-stage generative/contrastive/judge-curated training pipeline yields state-of-the-art MLLM-based retrieval accuracy at lower inference cost.
-
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
MOON2.0 combines modality-routed experts, intra-product image-text alignment, MLLM-generated data augmentation, and dynamic sample filtering to reach state-of-the-art zero-shot e-commerce product understanding.
-
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
Injecting an average-pooled visual embedding into every text token improves hallucination-benchmark scores of Video-LLaVA by small single-digit amounts.
-
Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
The abstract claims a new multi-rationale benchmark and a training-free contrastive framework for explainable object recognition, but the submitted full text is a different paper on M-valley twisted TMDs, so none of t...
-
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.
Discussion (0). Continue with ORCID to comment.