Pith. sign in

REVIEW 13 cited by

MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.14475 v1 pith:PWUPYNX5 submitted 2024-12-19 cs.CV cs.CL

classification cs.CVcs.CL
keywords datamegapairsmodelsperformanceretrievalmultimodalsynthesisdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70$\times$ more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our produced dataset, well-trained models, and data synthesis pipeline will be made publicly available to facilitate the future development of this field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

    cs.CV 2026-08 conditional novelty 7.0 of 10

    UniME-R1 uses a failure-aware adviser to diagnose embedding mistakes from initial retrieval results and then either reranks candidates or re-retrieves with a feedback-based query rewrite.

  2. Douyin Multimodal Embedding Model Technical Report

    cs.IR 2026-08 conditional novelty 6.0 of 10

    Latent typed reasoning plus cross-conditional reconstruction during training improves multimodal retrieval accuracy while keeping inference a standard dense bi-encoder, yielding 74.8 (2B) and 78.4 (9B) on MMEB-v2.

  3. VIG-RL: Learning to Search and Insert for Verified Image Grounding

    cs.IR 2026-07 conditional novelty 6.0 of 10

    An RL-trained ReAct agent learns joint search–selection–insertion of retrieved authentic images, setting SOTA on MRAMG-Bench Verified Image Grounding.

  4. FreeRet: MLLMs as Training-Free Retrievers

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    FreeRet enables pretrained MLLMs to act as training-free retrievers via semantically grounded embeddings and reasoning-based reranking, outperforming models trained on millions of pairs on MMEB benchmarks.

  5. VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A CoT-annotated dataset (VaccineRAG) plus segment-level GRPO (Partial-GRPO) improves multimodal large language models' ability to ignore harmful retrieved samples in retrieval-augmented generation tasks.

  6. M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation

    cs.IR 2025-08 conditional novelty 6.0 of 10

    M2IO-R1 uses GRPO reinforcement learning to train a 3B model that selects and places retrieved images into generated text answers, improving MRAMG quality on several benchmarks.

  7. MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage training recipe that converts causal VLMs into bidirectional multimodal embedding models, achieving SOTA on MMEB.

  8. mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A systematic empirical study finds that for multimodal RAG, EVA-CLIP retrieval, listwise LVLM reranking, and feeding only the top-ranked document works best, with a self-reflection agent adding further gains.

  9. Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

    cs.CV 2025-05 conditional novelty 6.0 of 10

    UNITE combines curated multimodal training data and a modality-masked contrastive loss to achieve strong retrieval performance across text, image, and video tasks.

  10. mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A multimodal embedding model trained on 560K GPT-4o-synthesized examples, covering 93 languages and 7 modality combinations, reaches top average scores on MMEB and XTD.

  11. MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

    cs.CV 2025-11 conditional novelty 5.0 of 10

    MOON2.0 combines modality-routed experts, intra-product image-text alignment, MLLM-generated data augmentation, and dynamic sample filtering to reach state-of-the-art zero-shot e-commerce product understanding.

  12. Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Video-XL-2 cuts long-video inference cost with chunked pre-filling and query-gated dense-or-sparse KV reloading, reporting half the FLOPs and a third less decoding memory at roughly equal benchmark scores.

  13. Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A structured survey of multimodal RAG systems, covering datasets, benchmarks, methods, and open challenges, with a public resource repo.

Pith tools