Pith. sign in

REVIEW 11 cited by

MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.06565 v2 pith:7WGTVI7K submitted 2024-06-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords benchmarksevaluationqueriesmixevalarenachatbotefficientexisting
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Evaluating large language models (LLMs) is challenging. Traditional ground-truth-based benchmarks fail to capture the comprehensiveness and nuance of real-world queries, while LLM-as-judge benchmarks suffer from grading biases and limited query quantity. Both of them may also become contaminated over time. User-facing evaluation, such as Chatbot Arena, provides reliable signals but is costly and slow. In this work, we propose MixEval, a new paradigm for establishing efficient, gold-standard LLM evaluation by strategically mixing off-the-shelf benchmarks. It bridges (1) comprehensive and well-distributed real-world user queries and (2) efficient and fairly-graded ground-truth-based benchmarks, by matching queries mined from the web with similar queries from existing benchmarks. Based on MixEval, we further build MixEval-Hard, which offers more room for model improvement. Our benchmarks' advantages lie in (1) a 0.96 model ranking correlation with Chatbot Arena arising from the highly impartial query distribution and grading mechanism, (2) fast, cheap, and reproducible execution (6% of the time and cost of MMLU), and (3) dynamic evaluation enabled by the rapid and stable data update pipeline. We provide extensive meta-evaluation and analysis for our and existing LLM benchmarks to deepen the community's understanding of LLM evaluation and guide future research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

    cs.CL 2025-05 conditional novelty 7.0 of 10

    SPOT shows that state-of-the-art AI models detect fewer than one in five known errors in full scientific papers, with precision below 7%.

  2. Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Shortcut neuron patching suppresses benchmark-contamination shortcuts in LLMs and yields evaluation scores that strongly correlate with the external MixEval benchmark.

  3. Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A fully automatic LLM evaluation framework where all evaluated models serve as judges for one another reaches 97% Spearman correlation with human preference rankings while keeping cost sub-quadratic.

  4. WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A 50-million-conversation synthetic chat dataset built from 54 open-weight models, plus an SFT mix that beats Tulu-3's mix with fewer samples.

  5. How to Select Datapoints for Efficient Human Evaluation of NLG Models?

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Selecting human-evaluation items by metric variance, metric consistency, output diversity, or IRT-based informativeness matches random-sampling ranking accuracy with roughly 70% of the annotation budget in WMT23 and SummEval.

  6. Tuning LLM Judge Design Decisions for 1/1000 of the Cost

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A multi-fidelity, multi-objective search finds cheap open-weight LLM judges that match or outperform prior judge designs on several benchmarks.

  7. Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change

    cs.CY 2025-01 conditional novelty 6.0 of 10

    Fine-tuned GPT-3.5-turbo agrees with expert climate coders on social media claims as often as the experts agree with each other (alpha=0.89), but the study's open-source benchmark is weakened by a flawed prompt and ra...

  8. Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Automatic LLM rankers align well with humans on broad leaderboards but degrade sharply when ranking close-performing models, and instance-level judge accuracy does not predict system-level bencher quality.

  9. C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    C2LEVA is a bilingual, multi-task LLM benchmark that combines passive test-data renewal with active data watermarking to reduce contamination risk, and ranks 15 models.

  10. Challenges in Trustworthy Human Evaluation of Chatbots

    cs.HC 2024-12 conditional novelty 6.0 of 10

    Open chatbot leaderboards like Chatbot Arena can have their model rankings moved by several positions with only 10% low-quality or adversarial votes.

  11. LLM Alignment as Retriever Optimization: An Information Retrieval Perspective

    cs.CL 2025-02 conditional novelty 5.0 of 10

    LARPO, an iterative preference optimization method that adapts information retrieval techniques such as listwise ranking losses, hard negatives, and candidate lists, is claimed to substantially improve LLM alignment o...

Pith tools