Pith. sign in

REVIEW 10 cited by

LLMJudge: LLMs for Relevance Judgments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08896 v1 pith:N4VDSXE7 submitted 2024-08-09 cs.IR

classification cs.IR
keywords llmsrelevancejudgmentssearchchallengedatallmjudgegenerate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The LLMJudge challenge is organized as part of the LLM4Eval workshop at SIGIR 2024. Test collections are essential for evaluating information retrieval (IR) systems. The evaluation and tuning of a search system is largely based on relevance labels, which indicate whether a document is useful for a specific search and user. However, collecting relevance judgments on a large scale is costly and resource-intensive. Consequently, typical experiments rely on third-party labelers who may not always produce accurate annotations. The LLMJudge challenge aims to explore an alternative approach by using LLMs to generate relevance judgments. Recent studies have shown that LLMs can generate reliable relevance judgments for search systems. However, it remains unclear which LLMs can match the accuracy of human labelers, which prompts are most effective, how fine-tuned open-source LLMs compare to closed-source LLMs like GPT-4, whether there are biases in synthetically generated data, and if data leakage affects the quality of generated labels. This challenge will investigate these questions, and the collected data will be released as a package to support automatic relevance judgment research in information retrieval and search.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look

    cs.IR 2024-11 conditional novelty 7.0 of 10

    Fully automatic LLM relevance judgments rank TREC 2024 retrieval runs as well as human judgments, and human-in-the-loop assistance provides no clear extra benefit.

  2. HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems

    cs.MA 2026-08 conditional novelty 6.0 of 10

    A hierarchical multi-agent LLM framework, in which a coordinating analyzer synthesizes query and item specialist outputs, beats flat, staged, and ensemble LLM relevance judges on five content search datasets.

  3. LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Large language models encode query-document relevance as a linearly decodable internal signal that strengthens in middle-to-late layers and, in several models, outperforms their own generated judgments.

  4. QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A fine-tuned ensemble of a generative small language model and an embedding model outperformed zero-shot LLMs on Korean search relevance labeling, with reported Cohen's kappa of 0.646 versus 0.387 and 60x lower latency.

  5. Benchmarking LLM-based Relevance Judgment Methods

    cs.IR 2025-04 conditional novelty 6.0 of 10

    On TREC DL 2019-2021 and ANTIQUE, pairwise LLM judgments align best with human labels, while binary and UMBRELA graded judgments agree most with human-based system rankings.

  6. AudioBoost: Increasing Audiobook Retrievability in Spotify Search with Synthetic Query Generation

    cs.IR 2025-09 conditional novelty 5.0 of 10

    LLM-generated synthetic queries, indexed in both autocomplete and document content, produced small but statistically significant gains in audiobook impressions, clicks, and exploratory searches in a production A/B test.

  7. Does UMBRELA Work on Other LLMs?

    cs.IR 2025-07 conditional novelty 5.0 of 10

    UMBRELA relevance judgments made with DeepSeek V3 are close to GPT-4o, and even small models preserve leaderboard rankings although per-document agreement with humans drops.

  8. Evaluating Contrastive Feedback for Effective User Simulations

    cs.IR 2025-05 conditional novelty 5.0 of 10

    Providing simulated search users with summaries of both relevant and irrelevant documents improves retrieval effectiveness on TREC Core17 but not on Core18, and no significance tests are reported.

  9. JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment

    cs.IR 2024-12 conditional novelty 5.0 of 10

    Ensembling small open-source LLMs as relevance judges achieves human-correlation scores competitive with GPT-4-based judges on the LLMJudge benchmark.

  10. Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

    cs.IR 2025-07 reject novelty 4.0 of 10

    LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.

Pith tools