REVIEW 6 cited by
State of What Art? A Call for Multi-Prompt LLM Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advances in large language models (LLMs) have led to the development of various evaluation benchmarks. These benchmarks typically rely on a single instruction template for evaluating all LLMs on a specific task. In this paper, we comprehensively analyze the brittleness of results obtained via single-prompt evaluations across 6.5M instances, involving 20 different LLMs and 39 tasks from 3 benchmarks. To improve robustness of the analysis, we propose to evaluate LLMs with a set of diverse prompts instead. We discuss tailored evaluation metrics for specific use cases (e.g., LLM developers vs. developers interested in a specific downstream task), ensuring a more reliable and meaningful assessment of LLM capabilities. We then implement these criteria and conduct evaluations of multiple models, providing insights into the true strengths and limitations of current LLMs.
Forward citations
Cited by 6 Pith papers
-
JuStRank: Benchmarking LLM Judges for System Ranking
JuStRank ranks AI judges by how well their aggregated scores reproduce the Chatbot Arena human system ranking, revealing that judge realization and bias, not just model size, determine ranking quality.
-
Evalita-LLM: Benchmarking Large Language Models on Italian
Evalita-LLM is a native-Italian, multi-prompt benchmark for LLMs built from ten Evalita datasets, with development-phase scores for six mid-size instruction-tuned models.
-
Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative
An LLM-powered storylets framework lets authors write natural-language triggers that fire at appropriate moments, supporting responsive interactive narratives with modest authoring effort.
-
LCTG Bench: LLM Controlled Text Generation Benchmark
LCTG Bench is a new Japanese benchmark that scores LLM controllability on format, character count, keyword, and prohibited-word constraints across summarization, ad text, and pros/cons generation.
-
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.
-
Personalizing Education through an Adaptive LMS with Integrated LLMs
The paper presents a hybrid expert-system and LLM adaptive LMS prototype and a benchmark of ten LLMs on standardized tests, showing self-hosted models are competitive with proprietary ones in reading, writing, and cod...
Discussion (0). Continue with ORCID to comment.