REVIEW 3 cited by
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty of LLM-as-a-Judge evaluations. While uncertainty quantification has been well-studied in other domains, applying it effectively to LLMs poses unique challenges due to their complex decision-making capabilities and computational demands. In this paper, we introduce a novel method for quantifying uncertainty designed to enhance the trustworthiness of LLM-as-a-Judge evaluations. The method quantifies uncertainty by analyzing the relationships between generated assessments and possible ratings. By cross-evaluating these relationships and constructing a confusion matrix based on token probabilities, the method derives labels of high or low uncertainty. We evaluate our method across multiple benchmarks, demonstrating a strong correlation between the accuracy of LLM evaluations and the derived uncertainty scores. Our findings suggest that this method can significantly improve the reliability and consistency of LLM-as-a-Judge evaluations.
Forward citations
Cited by 3 Pith papers
-
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
An IRT-based two-phase diagnostic framework with four metrics (CV, ρ, θratio, DW) for measuring LLM-judge intrinsic consistency and human alignment.
-
A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development
An LLM-based evaluator for agile epics was built from a new rubric and tested with 17 product managers, who found it useful but limited by lack of domain knowledge and rigid scoring.
-
A Survey of Calibration Process for Black-Box LLMs
A survey that organizes existing techniques for estimating and correcting confidence scores of black-box large language models into a two-step calibration process.
Discussion (0). Continue with ORCID to comment.