REVIEW 10 cited by
LLMJudge: LLMs for Relevance Judgments
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The LLMJudge challenge is organized as part of the LLM4Eval workshop at SIGIR 2024. Test collections are essential for evaluating information retrieval (IR) systems. The evaluation and tuning of a search system is largely based on relevance labels, which indicate whether a document is useful for a specific search and user. However, collecting relevance judgments on a large scale is costly and resource-intensive. Consequently, typical experiments rely on third-party labelers who may not always produce accurate annotations. The LLMJudge challenge aims to explore an alternative approach by using LLMs to generate relevance judgments. Recent studies have shown that LLMs can generate reliable relevance judgments for search systems. However, it remains unclear which LLMs can match the accuracy of human labelers, which prompts are most effective, how fine-tuned open-source LLMs compare to closed-source LLMs like GPT-4, whether there are biases in synthetically generated data, and if data leakage affects the quality of generated labels. This challenge will investigate these questions, and the collected data will be released as a package to support automatic relevance judgment research in information retrieval and search.
Forward citations
Cited by 10 Pith papers
-
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
Fully automatic LLM relevance judgments rank TREC 2024 retrieval runs as well as human judgments, and human-in-the-loop assistance provides no clear extra benefit.
-
HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems
A hierarchical multi-agent LLM framework, in which a coordinating analyzer synthesizes query and item specialist outputs, beats flat, staged, and ensemble LLM relevance judges on five content search datasets.
-
LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal
Large language models encode query-document relevance as a linearly decodable internal signal that strengthens in middle-to-late layers and, in several models, outperforms their own generated judgments.
-
QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines
A fine-tuned ensemble of a generative small language model and an embedding model outperformed zero-shot LLMs on Korean search relevance labeling, with reported Cohen's kappa of 0.646 versus 0.387 and 60x lower latency.
-
Benchmarking LLM-based Relevance Judgment Methods
On TREC DL 2019-2021 and ANTIQUE, pairwise LLM judgments align best with human labels, while binary and UMBRELA graded judgments agree most with human-based system rankings.
-
AudioBoost: Increasing Audiobook Retrievability in Spotify Search with Synthetic Query Generation
LLM-generated synthetic queries, indexed in both autocomplete and document content, produced small but statistically significant gains in audiobook impressions, clicks, and exploratory searches in a production A/B test.
-
Does UMBRELA Work on Other LLMs?
UMBRELA relevance judgments made with DeepSeek V3 are close to GPT-4o, and even small models preserve leaderboard rankings although per-document agreement with humans drops.
-
Evaluating Contrastive Feedback for Effective User Simulations
Providing simulated search users with summaries of both relevant and irrelevant documents improves retrieval effectiveness on TREC Core17 but not on Core18, and no significance tests are reported.
-
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment
Ensembling small open-source LLMs as relevance judges achieves human-correlation scores competitive with GPT-4-based judges on the LLMJudge benchmark.
-
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.
Discussion (0). Continue with ORCID to comment.