Pith. sign in

REVIEW 9 cited by

A Closer Look into Automatic Evaluation Using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05657 v1 pith:T6UXLZHD submitted 2023-10-09 cs.CL

classification cs.CL
keywords evaluationratingsg-evalhumanllmsdetailslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Using large language models (LLMs) to evaluate text quality has recently gained popularity. Some prior works explore the idea of using LLMs for evaluation, while they differ in some details of the evaluation process. In this paper, we analyze LLM evaluation (Chiang and Lee, 2023) and G-Eval (Liu et al., 2023), and we discuss how those details in the evaluation process change how well the ratings given by LLMs correlate with human ratings. We find that the auto Chain-of-Thought (CoT) used in G-Eval does not always make G-Eval more aligned with human ratings. We also show that forcing the LLM to output only a numeric rating, as in G-Eval, is suboptimal. Last, we reveal that asking the LLM to explain its own ratings consistently improves the correlation between the ChatGPT and human ratings and pushes state-of-the-art (SoTA) correlations on two meta-evaluation datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Collaborative Decoding fuses a knowledge-conditioned and a context-only token distribution with confidence- and divergence-based weights plus knowledge-aware reranking, improving faithfulness while keeping expressiven...

  2. FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new benchmark and 3D decomposition paradigm for factuality evaluation of interpretive claims about contact center conversations, with best LLM-judge F1 of 0.86.

  3. Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    PrefBERT, a 150M-parameter reward model trained on human quality ratings, outperforms ROUGE-L and BERTScore as a GRPO reward signal for open-ended long-form generation.

  4. An Empirical Study of Evaluating Long-form Question Answering

    cs.IR 2025-04 conditional novelty 5.0 of 10

    In long-form QA, LLM-based evaluators correlate with human judgments better than ROUGE or BERTScore, but they are biased by answer length, question type, self-reinforcement, and rare-word usage, and fine-grained promp...

  5. Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

    cs.CL 2025-04 conditional novelty 5.0 of 10

    A reference-based Likert scoring method with analyze-rate prompting improves LLM judges' agreement with human creativity rankings on the TTCW benchmark, though the headline result is partly fitted to the test set.

  6. Can Large Language Models Serve as Evaluators for Code Summarization?

    cs.SE 2024-12 conditional novelty 5.0 of 10

    An LLM prompt that makes the model role-play a code reviewer scores code summaries with 81.59% Spearman correlation with human judgment, outperforming BLEU and BERTScore on a 300-sample benchmark.

  7. Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources

    cs.CY 2025-01 conditional novelty 4.0 of 10

    Refining an LLM auto-evaluator with expert-teacher themes and few-shot examples improved agreement with human scores on quiz quality, but only on the same questions used for refinement.

  8. Do LLMs Agree on the Creativity Evaluation of Alternative Uses?

    cs.AI 2024-11 conditional novelty 4.0 of 10

    Four LLMs show high agreement when scoring and ranking alternative uses and do not favor their own outputs, but the accuracy benchmark is derived from the generation prompts rather than human judgment.

  9. SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions

    cs.AI 2024-11 reject novelty 4.0 of 10

    SRSA, a router that picks between direct, parallel, and planning searches for each query, improves informativeness and completeness on a new contextual-query benchmark while using fewer LLM inference steps than a ReAct agent.

Pith tools