Pith. sign in

REVIEW 11 cited by

Human-like Summarization Evaluation with ChatGPT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.02554 v1 pith:LYO2NW74 submitted 2023-04-05 cs.CL

classification cs.CL
keywords evaluationchatgptsummarizationdatasetshumanhuman-likemetricsability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating text summarization is a challenging problem, and existing evaluation metrics are far from satisfactory. In this study, we explored ChatGPT's ability to perform human-like summarization evaluation using four human evaluation methods on five datasets. We found that ChatGPT was able to complete annotations relatively smoothly using Likert scale scoring, pairwise comparison, Pyramid, and binary factuality evaluation. Additionally, it outperformed commonly used automatic evaluation metrics on some datasets. Furthermore, we discussed the impact of different prompts, compared its performance with that of human evaluation, and analyzed the generated explanations and invalid responses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 38 citations worldwide. Full citation record

  1. Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

    cs.AI 2026-01 conditional novelty 6.0 of 10

    An IRT-based two-phase diagnostic framework with four metrics (CV, ρ, θratio, DW) for measuring LLM-judge intrinsic consistency and human alignment.

  2. Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning

    cs.AI 2025-09 reject novelty 6.0 of 10

    A two-stage pattern-aware tool-integrated reasoning method raises code usage and code-plus-correct metrics on math benchmarks, but the paper conflates Code@1 with problem-solving accuracy in its headline claims.

  3. Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A regression controlling for human-rated quality detects positive self- and family-bias in several LLM judges, including GPT-4o and Claude 3.5 Sonnet.

  4. From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    ProxyReward trains long-form generation models by rewarding how well an AI judge can answer generated yes/no questions about the response, improving open-source models on ProxyQA.

  5. From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?

    cs.AI 2025-12 conditional novelty 5.0 of 10

    The paper proposes a judgment-and-steering mediation framework for LLMs in online flame wars and reports that API-based models outperform open-source ones on a principle-based LLM-judged evaluation, with the toxicity-...

  6. AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    AraHalluEval introduces a 12-indicator Arabic hallucination taxonomy and finds factual errors dominate, with Allam competitive against reasoning models.

  7. Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A mixed-methods framework that elicits stakeholder preferences and steers LLMs to generate tailored summaries of health simulations, presented without empirical validation.

  8. User Perceptions of an LLM-Based Chatbot for Cognitive Reappraisal of Stress: Feasibility Study

    cs.HC 2026-01 conditional novelty 4.0 of 10

    A GPT-4o chatbot guiding employees through an 11-step reappraisal script was associated with small short-term reductions in self-reported stress and improved stress mindset in an uncontrolled feasibility study.

  9. Byzantine-Robust Decentralized Coordination of LLM Agents

    cs.DC 2025-07 conditional novelty 4.0 of 10

    A leaderless, Byzantine-robust LLM agent coordination protocol selects answers by geometric-median aggregation of evaluator scores instead of leader-based quorum voting.

  10. Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

    cs.IR 2025-07 reject novelty 4.0 of 10

    LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.

  11. Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance

    cs.LG 2025-06 reject novelty 4.0 of 10

    A frozen 7B-8B LLM with a JSON rubric and a small LoRA adapter is claimed to outperform 27B-70B reward models and enable 92% GSM-8K exact match under online PPO.

Pith tools