Pith. sign in

REVIEW 3 cited by

Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13439 v2 pith:RV2YWKCS submitted 2024-06-19 cs.CL

classification cs.CL
keywords llmsevaluatoraccuracyanswerscurrentdropsevaluationsevaluators
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are increasingly relied upon to evaluate text outputs of other LLMs, thereby influencing leaderboards and development decisions. However, concerns persist over the accuracy of these assessments and the potential for misleading conclusions. In this work, we investigate the effectiveness of LLMs as evaluators for text generation tasks. We propose FBI, a novel framework designed to examine the proficiency of Evaluator LLMs in assessing four critical abilities in other LLMs: factual accuracy, instruction following, coherence in long-form writing, and reasoning proficiency. By introducing targeted perturbations in answers generated by LLMs, that clearly impact one of these key capabilities, we test whether an Evaluator LLM can detect these quality drops. By creating a total of 2400 perturbed answers covering 22 perturbation categories, we conduct a comprehensive study using different evaluation strategies on five prominent LLMs commonly used as evaluators in the literature. Our findings reveal significant shortcomings in current Evaluator LLMs, which failed to identify quality drops in over 50\% of cases on average. Single-answer and pairwise evaluations demonstrated notable limitations, whereas reference-based evaluations showed comparatively better performance. These results underscore the unreliable nature of current Evaluator LLMs and advocate for cautious implementation in practical applications. Code and data are available at https://github.com/AI4Bharat/FBI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tuning LLM Judge Design Decisions for 1/1000 of the Cost

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A multi-fidelity, multi-objective search finds cheap open-weight LLM judges that match or outperform prior judge designs on several benchmarks.

  2. An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The reliability of LLM-as-a-Judge depends strongly on scoring rubrics and reference answers; sampling with averaging outperforms greedy decoding, and chain-of-thought reasoning adds little when rubrics are clear.

  3. OpenCoderRank: Personalized Technical Assessments with Generative AI

    cs.SE 2025-09 unverdicted novelty 4.0 of 10

    OpenCoderRank provides a self-hosted, customizable system for time-bound technical assessments with automatic grading via BERTScore and LLM evaluation.

Pith tools