Pith. sign in

REVIEW 15 cited by

Evaluating Large Language Models: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19736 v3 pith:RGXH2KSS submitted 2023-10-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationllmscomprehensivepotentialalignmentbeencapabilitiesdevelopment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have demonstrated remarkable capabilities across a broad spectrum of tasks. They have attracted significant attention and been deployed in numerous downstream applications. Nevertheless, akin to a double-edged sword, LLMs also present potential risks. They could suffer from private data leaks or yield inappropriate, harmful, or misleading content. Additionally, the rapid progress of LLMs raises concerns about the potential emergence of superintelligent systems without adequate safeguards. To effectively capitalize on LLM capacities as well as ensure their safe and beneficial development, it is critical to conduct a rigorous and comprehensive evaluation of LLMs. This survey endeavors to offer a panoramic perspective on the evaluation of LLMs. We categorize the evaluation of LLMs into three major groups: knowledge and capability evaluation, alignment evaluation and safety evaluation. In addition to the comprehensive review on the evaluation methodologies and benchmarks on these three aspects, we collate a compendium of evaluations pertaining to LLMs' performance in specialized domains, and discuss the construction of comprehensive evaluation platforms that cover LLM evaluations on capabilities, alignment, safety, and applicability. We hope that this comprehensive overview will stimulate further research interests in the evaluation of LLMs, with the ultimate goal of making evaluation serve as a cornerstone in guiding the responsible development of LLMs. We envision that this will channel their evolution into a direction that maximizes societal benefit while minimizing potential risks. A curated list of related papers has been publicly available at https://github.com/tjunlp-lab/Awesome-LLMs-Evaluation-Papers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A clinician-validated benchmark for patient-facing health AI agents shows that even frontier models fail triage in up to a quarter of realistic tool-using conversations.

  2. Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Decoupling LLM-agent planning from summarization and rewarding tool-call completeness rather than final-answer correctness improves planning by 8-12% and end-to-end answers by 5-6% over end-to-end RL baselines.

  3. SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs

    cs.CL 2025-07 conditional novelty 6.0 of 10

    SKA-Bench is a fine-grained QA benchmark across KG, table, and hybrid formats that shows current LLMs remain sensitive to noise and order and often hallucinate instead of rejecting unanswerable inputs.

  4. Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Shortcut neuron patching suppresses benchmark-contamination shortcuts in LLMs and yields evaluation scores that strongly correlate with the external MixEval benchmark.

  5. Human-Centric Evaluation for Foundation Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    In free-form research collaborations rated by humans, Grok 3 scores highest, followed by DeepSeek R1 and Gemini 2.5; OpenAI o3 mini trails.

  6. Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLMs align more closely with human ratings on affective and cognitive word norms than on sensory-perceptual norms, suggesting a gap tied to embodied experience.

  7. From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Fine-tuning GPT-3.5 and Llama 2 on r/Anxiety posts improves readability but raises toxicity and bias while reducing empathy and reflection.

  8. Efficient Sequential Evaluation of Large Language Models

    stat.ML 2026-07 conditional novelty 5.0 of 10

    A confidence-sequence framework for sequentially estimating an LLM's average benchmark accuracy under adaptive question selection, with growth-oriented sampling rules that in practice often lose to uniform sampling.

  9. User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors introduce an entropy-based framework that uses user behavior prediction as a measure of LLM generalization, and find GPT-4o outperforms GPT-4o-mini and Llama-3.1 on movie and music recommendation tasks.

  10. Benchmarking the Pedagogical Knowledge of Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors release an open benchmark of 1,143 pedagogical knowledge questions from Chilean teacher exams and report accuracy, cost, and size trade-offs for 97 large language models.

  11. Cognitive Agents Powered by Large Language Models for Agile Software Project Management

    cs.SE 2025-08 reject novelty 4.0 of 10

    LLM agents acting as Agile roles produced plausible project artifacts in simulation, but the claimed improvements over human teams are unsupported because no comparison or validated metrics are provided.

  12. OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Using Olmo to extract atomic facts and Gemma to verify them against Wikipedia, OpenFActScore reproduces the original FActScore ranking of 10 LLMs with a Pearson correlation above 0.99.

  13. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  14. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

  15. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

Pith tools