REVIEW 6 cited by
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Copious amounts of relevance judgments are necessary for the effective training and accurate evaluation of retrieval systems. Conventionally, these judgments are made by human assessors, rendering this process expensive and laborious. A recent study by Thomas et al. from Microsoft Bing suggested that large language models (LLMs) can accurately perform the relevance assessment task and provide human-quality judgments, but unfortunately their study did not yield any reusable software artifacts. Our work presents UMBRELA (a recursive acronym that stands for UMbrela is the Bing RELevance Assessor), an open-source toolkit that reproduces the results of Thomas et al. using OpenAI's GPT-4o model and adds more nuance to the original paper. Across Deep Learning Tracks from TREC 2019 to 2023, we find that LLM-derived relevance judgments correlate highly with rankings generated by effective multi-stage retrieval systems. Our toolkit is designed to be easily extensible and can be integrated into existing multi-stage retrieval and evaluation pipelines, offering researchers a valuable resource for studying retrieval evaluation methodologies. UMBRELA will be used in the TREC 2024 RAG Track to aid in relevance assessments, and we envision our toolkit becoming a foundation for further innovation in the field. UMBRELA is available at https://github.com/castorini/umbrela.
Forward citations
Cited by 6 Pith papers
-
mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health
The authors introduce and release two new benchmarks for maternal-health RAG evaluation, built from expert sources with graded labels and disclosed limitations rather than binary judgments or new question authoring.
-
LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal
Large language models encode query-document relevance as a linearly decodable internal signal that strengthens in middle-to-late layers and, in several models, outperforms their own generated judgments.
-
Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
Structured schema.org metadata still gives dataset-retrieval agents a large precision advantage for machine-actionable data.
-
Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search
A production VLM-based relevance-labeling pipeline at Pinterest search produces human-aligned sDCG@K metrics and about a 6× smaller minimum detectable effect in A/B tests.
-
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
A hotel retrieval system combining small query encoders with large document encoders, multi-task training, and pooled image features outperforms prior multimodal retrievers on four synthetic test sets.
-
Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems
The paper adds Type II error metrics to the evaluation of relevance judgment sets and shows that balanced accuracy and Matthews correlation can summarize qrels' discriminative power in one number.
Discussion (0). Continue with ORCID to comment.