Pith. sign in

REVIEW 6 cited by

UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06519 v1 pith:GEBUETWA submitted 2024-06-10 cs.IR

classification cs.IR
keywords umbrelarelevancejudgmentsretrievalbingevaluationtoolkitassessor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Copious amounts of relevance judgments are necessary for the effective training and accurate evaluation of retrieval systems. Conventionally, these judgments are made by human assessors, rendering this process expensive and laborious. A recent study by Thomas et al. from Microsoft Bing suggested that large language models (LLMs) can accurately perform the relevance assessment task and provide human-quality judgments, but unfortunately their study did not yield any reusable software artifacts. Our work presents UMBRELA (a recursive acronym that stands for UMbrela is the Bing RELevance Assessor), an open-source toolkit that reproduces the results of Thomas et al. using OpenAI's GPT-4o model and adds more nuance to the original paper. Across Deep Learning Tracks from TREC 2019 to 2023, we find that LLM-derived relevance judgments correlate highly with rankings generated by effective multi-stage retrieval systems. Our toolkit is designed to be easily extensible and can be integrated into existing multi-stage retrieval and evaluation pipelines, offering researchers a valuable resource for studying retrieval evaluation methodologies. UMBRELA will be used in the TREC 2024 RAG Track to aid in relevance assessments, and we envision our toolkit becoming a foundation for further innovation in the field. UMBRELA is available at https://github.com/castorini/umbrela.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    The authors introduce and release two new benchmarks for maternal-health RAG evaluation, built from expert sources with graded labels and disclosed limitations rather than binary judgments or new question authoring.

  2. LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Large language models encode query-document relevance as a linearly decodable internal signal that strengthens in middle-to-late layers and, in several models, outperforms their own generated judgments.

  3. Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval

    cs.IR 2026-05 unverdicted novelty 6.0 of 10

    Structured schema.org metadata still gives dataset-retrieval agents a large precision advantage for machine-actionable data.

  4. Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

    cs.IR 2026-08 conditional novelty 5.0 of 10

    A production VLM-based relevance-labeling pipeline at Pinterest search produces human-aligned sDCG@K metrics and about a 6× smaller minimum detectable effect in A/B tests.

  5. HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval

    cs.IR 2025-06 conditional novelty 5.0 of 10

    A hotel retrieval system combining small query encoders with large document encoders, multi-task training, and pooled image features outperforms prior multimodal retrievers on four synthetic test sets.

  6. Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems

    cs.IR 2025-07 conditional novelty 4.0 of 10

    The paper adds Type II error metrics to the evaluation of relevance judgment sets and shows that balanced accuracy and Matthews correlation can summarize qrels' discriminative power in one number.

Pith tools