Pith. sign in

REVIEW 3 cited by

Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19740 v2 pith:3ALB7XMR submitted 2023-10-30 cs.CL

classification cs.CL
keywords criteriaevaluatorsevaluationtaskshumanlanguagealignmentdiverse
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Previous work adopts large language models (LLMs) as evaluators to evaluate natural language process (NLP) tasks. However, certain shortcomings, e.g., fairness, scope, and accuracy, persist for current LLM evaluators. To analyze whether LLMs can serve as reliable alternatives to humans, we examine the fine-grained alignment between LLM evaluators and human annotators, particularly in understanding the target evaluation tasks and conducting evaluations that meet diverse criteria. This paper explores both conventional tasks (e.g., story generation) and alignment tasks (e.g., math reasoning), each with different evaluation criteria. Our analysis shows that 1) LLM evaluators can generate unnecessary criteria or omit crucial criteria, resulting in a slight deviation from the experts. 2) LLM evaluators excel in general criteria, such as fluency, but face challenges with complex criteria, such as numerical reasoning. We also find that LLM-pre-drafting before human evaluation can help reduce the impact of human subjectivity and minimize annotation outliers in pure human evaluation, leading to more objective evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    ExPerT is a reference-based LLM evaluation metric that extracts and matches atomic aspects, scores content and style, and reports 0.74 human alignment on LongLaMP, a 7.2% relative gain over GEMBA and G-Eval.

  2. Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Evaluation Agent is an LLM-agent framework that evaluates visual generative models with a handful of samples per round, claiming a 10x time reduction while keeping conclusions within one tier of full-benchmark results...

  3. Do LLMs Agree on the Creativity Evaluation of Alternative Uses?

    cs.AI 2024-11 conditional novelty 4.0 of 10

    Four LLMs show high agreement when scoring and ranking alternative uses and do not favor their own outputs, but the accuracy benchmark is derived from the generation prompts rather than human judgment.

Pith tools