REVIEW 13 cited by
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70$\times$ more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our produced dataset, well-trained models, and data synthesis pipeline will be made publicly available to facilitate the future development of this field.
Forward citations
Cited by 13 Pith papers
-
Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
UniME-R1 uses a failure-aware adviser to diagnose embedding mistakes from initial retrieval results and then either reranks candidates or re-retrieves with a feedback-based query rewrite.
-
Douyin Multimodal Embedding Model Technical Report
Latent typed reasoning plus cross-conditional reconstruction during training improves multimodal retrieval accuracy while keeping inference a standard dense bi-encoder, yielding 74.8 (2B) and 78.4 (9B) on MMEB-v2.
-
VIG-RL: Learning to Search and Insert for Verified Image Grounding
An RL-trained ReAct agent learns joint search–selection–insertion of retrieved authentic images, setting SOTA on MRAMG-Bench Verified Image Grounding.
-
FreeRet: MLLMs as Training-Free Retrievers
FreeRet enables pretrained MLLMs to act as training-free retrievers via semantically grounded embeddings and reasoning-based reranking, outperforming models trained on millions of pairs on MMEB benchmarks.
-
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
A CoT-annotated dataset (VaccineRAG) plus segment-level GRPO (Partial-GRPO) improves multimodal large language models' ability to ignore harmful retrieved samples in retrieval-augmented generation tasks.
-
M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation
M2IO-R1 uses GRPO reinforcement learning to train a 3B model that selects and places retrieved images into generated text answers, improving MRAMG quality on several benchmarks.
-
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
A two-stage training recipe that converts causal VLMs into bidirectional multimodal embedding models, achieving SOTA on MMEB.
-
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
A systematic empirical study finds that for multimodal RAG, EVA-CLIP retrieval, listwise LVLM reranking, and feeding only the top-ranked document works best, with a self-reflection agent adding further gains.
-
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
UNITE combines curated multimodal training data and a modality-masked contrastive loss to achieve strong retrieval performance across text, image, and video tasks.
-
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
A multimodal embedding model trained on 560K GPT-4o-synthesized examples, covering 93 languages and 7 modality combinations, reaches top average scores on MMEB and XTD.
-
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
MOON2.0 combines modality-routed experts, intra-product image-text alignment, MLLM-generated data augmentation, and dynamic sample filtering to reach state-of-the-art zero-shot e-commerce product understanding.
-
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Video-XL-2 cuts long-video inference cost with chunked pre-filling and query-gated dense-or-sparse KV reloading, reporting half the FLOPs and a third less decoding memory at roughly equal benchmark scores.
-
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
A structured survey of multimodal RAG systems, covering datasets, benchmarks, methods, and open challenges, with a public resource repo.
Discussion (0). Continue with ORCID to comment.