Pith. sign in

REVIEW 4 cited by

MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13716 v2 pith:UQZDXALH submitted 2024-10-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords judgemirage-benchbenchmarkevaluationgenerationmultilingualsystemsarena-based
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Traditional retrieval-augmented generation (RAG) benchmarks evaluate systems using heuristic-based metrics, but these require human preferences as the ground truth for reference. In contrast, arena-based benchmarks, where systems compete against each other, require an expensive large language model (LLM) as a judge for a reliable evaluation. We present a simple efficient technique to combine the best of both worlds. The idea is to train a surrogate judge using heuristic metrics as input, to output the LLM as a judge prediction. In our work, we develop MIRAGE-Bench, a synthetic arena-based RAG benchmark for 18 diverse languages on Wikipedia focused on multilingual answer generation evaluation. It extensively couples both heuristic features and LLM as a judge for evaluation. We benchmark 19 multilingual LLMs, and observe a high correlation (Kendall Tau ($\tau$) = 0.909) using our surrogate judge and between GPT-4o as a teacher using the Bradley-Terry framework. Our results show proprietary and large open-source LLMs currently dominate on MIRAGE-Bench. Our code and datasets are made publicly available here: https://github.com/vectara/mirage-bench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. XRAG: Cross-lingual Retrieval-Augmented Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    XRAG is a new cross-lingual RAG benchmark showing that large language models frequently ignore the question language and fail at combining evidence from documents in different languages.

  2. The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models

    cs.IR 2025-04 conditional novelty 6.0 of 10

    A fully automatic LLM-based nugget evaluation for RAG systems matches human assessments at the run level on TREC 2024, with stronger agreement when only nugget assignment is automated.

  3. Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses

    cs.IR 2025-04 conditional novelty 5.0 of 10

    Nugget-based fact-recall scores on Search Arena battles correlate with human preferences, with about 54% agreement, and offer finer-grained diagnostic signals than raw win/loss votes.

  4. Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets

    cs.IR 2025-04 conditional novelty 3.0 of 10

    A systematic review of 63 RAG evaluation papers concludes that LLM-based automation is feasible across dataset generation, retrieval scoring, and answer evaluation, but only six studies directly compare LLM judges wit...

Pith tools