Pith. sign in

REVIEW 7 cited by

Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16365 v4 pith:BYSED45S submitted 2024-11-25 cs.CL

classification cs.CL
keywords multi-modalmetricsmodelsdatafoundationmodelaugmentedcomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We present a systematic investigation of Multi-modal Retrieval Augmented Multi-modal Generation (M$^2$RAG), a novel task that enables foundation models to process multi-modal web content and generate multi-modal responses, which exhibits better information density and readability. Despite its potential impact, M$^2$RAG remains understudied, lacking comprehensive analysis and high-quality data resources. To address this gap, we establish a comprehensive benchmark through a rigorous data curation pipeline, and employ text-modal metrics and multi-modal metrics based on foundation models for evaluation. We further propose several strategies for foundation models to process M$^2$RAG task effectively and construct a training set by filtering high-quality samples using our designed metrics. Our extensive experiments demonstrate the reliability of our proposed metrics, a landscape of model performance within our designed strategies, and show that our fine-tuned 7B-8B models outperform the GPT-4o model and approach the state-of-the-art OpenAI o3-mini. Additionally, we perform fine-grained analyses across diverse domains and validate the effectiveness of our designs in data curation pipeline. All resources, including codes, datasets, and model weights, will be publicly released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VIG-RL: Learning to Search and Insert for Verified Image Grounding

    cs.IR 2026-07 conditional novelty 6.0 of 10

    An RL-trained ReAct agent learns joint search–selection–insertion of retrieved authentic images, setting SOTA on MRAMG-Bench Verified Image Grounding.

  2. M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation

    cs.IR 2025-08 conditional novelty 6.0 of 10

    M2IO-R1 uses GRPO reinforcement learning to train a 3B model that selects and places retrieved images into generated text answers, improving MRAMG quality on several benchmarks.

  3. MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MVFT-JI trains a Q-Former VLM with two MLLM-generated retrieval tasks and fuses VLM and MLLM similarities at inference, achieving state-of-the-art zero-shot composed image retrieval on three benchmarks.

  4. MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A human-annotated benchmark with 4,800 QA pairs evaluates how well AI models can retrieve and generate interleaved text-and-image answers.

  5. Caption Injection for Optimization in Generative Search Engine

    cs.IR 2025-11 conditional novelty 5.0 of 10

    Adding VLM-generated, LLM-refined image captions into source text improves source visibility in generative search by about 1–2% relative, per the paper's G-Eval measurements on MRAMG.

  6. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  7. Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic Retrieval-Augmented Generation for Industry Challenges

    cs.AI 2025-06 unverdicted novelty 3.0 of 10

    A review that classifies Reasoning Agentic RAG into predefined (System 1-like) and agentic (System 2-like) workflows, surveying their designs and training strategies.

Pith tools