Pith. sign in

REVIEW 6 cited by

State of What Art? A Call for Multi-Prompt LLM Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00595 v3 pith:WLK5QSX5 submitted 2023-12-31 cs.CL

classification cs.CL
keywords llmsbenchmarksevaluationspecificdevelopersevaluationsmodelstask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in large language models (LLMs) have led to the development of various evaluation benchmarks. These benchmarks typically rely on a single instruction template for evaluating all LLMs on a specific task. In this paper, we comprehensively analyze the brittleness of results obtained via single-prompt evaluations across 6.5M instances, involving 20 different LLMs and 39 tasks from 3 benchmarks. To improve robustness of the analysis, we propose to evaluate LLMs with a set of diverse prompts instead. We discuss tailored evaluation metrics for specific use cases (e.g., LLM developers vs. developers interested in a specific downstream task), ensuring a more reliable and meaningful assessment of LLM capabilities. We then implement these criteria and conduct evaluations of multiple models, providing insights into the true strengths and limitations of current LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JuStRank: Benchmarking LLM Judges for System Ranking

    cs.CL 2024-12 conditional novelty 7.0 of 10

    JuStRank ranks AI judges by how well their aggregated scores reproduce the Chatbot Arena human system ranking, revealing that judge realization and bias, not just model size, determine ranking quality.

  2. Evalita-LLM: Benchmarking Large Language Models on Italian

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Evalita-LLM is a native-Italian, multi-prompt benchmark for LLMs built from ten Evalita datasets, with development-phase scores for six mid-size instruction-tuned models.

  3. Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative

    cs.HC 2025-01 conditional novelty 6.0 of 10

    An LLM-powered storylets framework lets authors write natural-language triggers that fire at appropriate moments, supporting responsive interactive narratives with modest authoring effort.

  4. LCTG Bench: LLM Controlled Text Generation Benchmark

    cs.CL 2025-01 conditional novelty 5.0 of 10

    LCTG Bench is a new Japanese benchmark that scores LLM controllability on format, character count, keyword, and prohibited-word constraints across summarization, ad text, and pros/cons generation.

  5. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.

  6. Personalizing Education through an Adaptive LMS with Integrated LLMs

    cs.AI 2025-01 conditional novelty 3.0 of 10

    The paper presents a hybrid expert-system and LLM adaptive LMS prototype and a benchmark of ten LLMs on standardized tests, showing self-hosted models are competitive with proprietary ones in reading, writing, and cod...

Pith tools