REVIEW 10 cited by
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibility and unreliability, leading users to prefer human voting platforms like Chatbot Arena. However, human evaluations require significant manual effort. To address this, we propose the Auto-Arena, an innovative framework that automates the entire evaluation process using LLM-powered agents. Firstly, an LLM examiner generates questions. Then, two LLM candidates engage in a multi-round peer battle based on individual questions, aiming at revealing their true performance differences. Finally, a committee of LLM judges collaboratively discusses and decides the winner, reducing bias and enhancing fairness. During the peer battles, we observe intriguing scenarios where the LLM candidates display competitive behaviors and even learn from the opponents. In our extensive experiments involving 15 recent LLMs, Auto-Arena shows a 92.14% correlation with human preferences, surpassing all previous expert-annotated benchmarks without any manual efforts. As a result, Auto-Arena offers a promising alternative to current human evaluation platforms for evaluating LLMs automatically.
Forward citations
Cited by 10 Pith papers
-
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
A fully automatic LLM evaluation framework where all evaluated models serve as judges for one another reaches 97% Spearman correlation with human preference rankings while keeping cost sub-quadratic.
-
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
A three-stage benchmark consisting of 982 pointing tasks, a live pairwise arena with 4,500 votes, and a robot manipulation study shows that pointing-supervised open models such as Molmo-72B can match proprietary model...
-
WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models
A code LLM fine-tuned on winner responses from pairwise expert battles, with instructions mined from chat templates, beats same-size baselines without proprietary LLMs.
-
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
An automated arena benchmark that simulates real users asking open-ended video questions, uses GPT-4o as judge, and ranks 11 large multimodal models via ELO ratings.
-
The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz
A new benchmark of 675 unsolvable questions finds that leading LLMs often fail to admit ignorance, scoring 62-68% even when 'I don't know' is the only correct choice.
-
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
An LLM-driven gate between robot planning and execution labels plans accept, reject, or escalate, reporting 81 percent accuracy and no direct accept/reject errors on small test sets.
-
Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
A debate-based evaluation protocol on 50 MMLU-Pro questions: fine-tuning on the test set boosts standard accuracy from 50% to 82% but not debate win rates.
-
ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data
A fine-tuned LLaMA-2-7B model trained with selected LLM-generated data improves extraction of chemical synthesis actions from experimental text.
-
ReviewInstruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models
A review-driven multi-agent pipeline turns single-turn instruction data into harder, more diverse multi-turn dialogues and improves a Llama2-13B model on MT-Bench and MMLU-Pro.
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
Discussion (0). Continue with ORCID to comment.