Pith. sign in

REVIEW 10 cited by

Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20267 v4 pith:PXZITQHJ submitted 2024-05-30 cs.CL

classification cs.CL
keywords auto-arenahumanevaluationllmspeerbattlesbenchmarkscandidates
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibility and unreliability, leading users to prefer human voting platforms like Chatbot Arena. However, human evaluations require significant manual effort. To address this, we propose the Auto-Arena, an innovative framework that automates the entire evaluation process using LLM-powered agents. Firstly, an LLM examiner generates questions. Then, two LLM candidates engage in a multi-round peer battle based on individual questions, aiming at revealing their true performance differences. Finally, a committee of LLM judges collaboratively discusses and decides the winner, reducing bias and enhancing fairness. During the peer battles, we observe intriguing scenarios where the LLM candidates display competitive behaviors and even learn from the opponents. In our extensive experiments involving 15 recent LLMs, Auto-Arena shows a 92.14% correlation with human preferences, surpassing all previous expert-annotated benchmarks without any manual efforts. As a result, Auto-Arena offers a promising alternative to current human evaluation platforms for evaluating LLMs automatically.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A fully automatic LLM evaluation framework where all evaluated models serve as judges for one another reaches 97% Spearman correlation with human preference rankings while keeping cost sub-quadratic.

  2. PointArena: Probing Multimodal Grounding Through Language-Guided Pointing

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A three-stage benchmark consisting of 982 pointing tasks, a live pairwise arena with 4,500 votes, and a robot manipulation study shows that pointing-supervised open models such as Molmo-72B can match proprietary model...

  3. WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A code LLM fine-tuned on winner responses from pairwise expert battles, with instructions mined from chat templates, beats same-size baselines without proprietary LLMs.

  4. VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An automated arena benchmark that simulates real users asking open-ended video questions, uses GPT-4o as judge, and ranks 11 large multimodal models via ELO ratings.

  5. The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz

    cs.CL 2024-11 reject novelty 6.0 of 10

    A new benchmark of 675 unsolvable questions finds that leading LLMs often fail to admit ignorance, scoring 62-68% even when 'I don't know' is the only correct choice.

  6. Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

    cs.RO 2026-08 conditional novelty 5.0 of 10

    An LLM-driven gate between robot planning and execution labels plans accept, reject, or escalate, reporting 81 percent accuracy and no direct accept/reject errors on small test sets.

  7. Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A debate-based evaluation protocol on 50 MMLU-Pro questions: fine-tuning on the test set boosts standard accuracy from 50% to 82% but not debate win rates.

  8. ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A fine-tuned LLaMA-2-7B model trained with selected LLM-generated data improves extraction of chemical synthesis actions from experimental text.

  9. ReviewInstruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A review-driven multi-agent pipeline turns single-turn instruction data into harder, more diverse multi-turn dialogues and improves a Llama2-13B model on MT-Bench and MMLU-Pro.

  10. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

Pith tools