League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models

Baosheng Wang, Enze Wang, Kai Chen, Qianhong Guo, Shuoyoucheng Ma, Tian Xia, Wei Xie, Xiaobing Sun, Xiaofang Cai, Xiaofeng Wang

Authors on Pith no claims yet

classification 💻 cs.AI cs.CL

keywords evaluationllmsleaguemodelsbenchmark-freecapabilitieslanguagelarge

0 comments

read the original abstract

Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination, opaque operation, and subjective preferences. To address these issues, we propose League of LLMs (LOL), a novel benchmark-free evaluation paradigm that organizes multiple LLMs into a self-governed league for multi-round mutual evaluation. LOL integrates four core criteria (dynamic, transparent, objective, and professional) to mitigate key limitations of existing paradigms. Experiments on eight mainstream LLMs in mathematics and programming demonstrate that LOL can effectively distinguish LLM capabilities while maintaining high internal ranking stability (Top-$k$ consistency $= 70.7\%$). Beyond ranking, LOL reveals empirical findings that are difficult for traditional paradigms to capture. For instance, ``memorization-based answering'' behaviors are observed in some models, and higher in-family scores are found in the OpenAI model family ($\Delta = 9$, $p < 0.05$). Finally, we make our framework and code publicly available as a valuable complement to the current LLM evaluation ecosystem.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

DeGenTWeb: A First Look at LLM-dominant Websites
cs.NI 2026-04 unverdicted novelty 5.0

DeGenTWeb shows LLM-dominant websites are common and increasing in Common Crawl and Bing search results, but accurate detection is getting harder with newer models.