Pith. sign in

REVIEW 14 cited by

Don't Make Your LLM an Evaluation Benchmark Cheater

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.01964 v1 pith:PNHPTCMO submitted 2023-11-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluationbenchmarksmodelbenchmarkllmsappropriateconcernsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models~(LLMs) have greatly advanced the frontiers of artificial intelligence, attaining remarkable improvement in model capacity. To assess the model performance, a typical approach is to construct evaluation benchmarks for measuring the ability level of LLMs in different aspects. Despite that a number of high-quality benchmarks have been released, the concerns about the appropriate use of these benchmarks and the fair comparison of different models are increasingly growing. Considering these concerns, in this paper, we discuss the potential risk and impact of inappropriately using evaluation benchmarks and misleadingly interpreting the evaluation results. Specially, we focus on a special issue that would lead to inappropriate evaluation, \ie \emph{benchmark leakage}, referring that the data related to evaluation sets is occasionally used for model training. This phenomenon now becomes more common since pre-training data is often prepared ahead of model test. We conduct extensive experiments to study the effect of benchmark leverage, and find that it can dramatically boost the evaluation results, which would finally lead to an unreliable assessment of model performance. To improve the use of existing evaluation benchmarks, we finally present several guidelines for both LLM developers and benchmark maintainers. We hope this work can draw attention to appropriate training and evaluation of LLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 18 citations worldwide. Full citation record

  1. Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps

    cs.NI 2025-10 conditional novelty 7.0 of 10

    Commercial AI video chat apps differ by 4× in video bitrate, 10× in framerate, and from zero to 10+ minutes of visual memory, with none replying in under 1.5 seconds.

  2. Can Vision Language Models Understand Mimed Actions?

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Vision-language models identify real actions with context far better than they identify mimed actions performed by 3D avatars, while humans are equally accurate on both.

  3. From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

    cs.CL 2026-07 conditional novelty 6.5 of 10

    On 2,520 programming tasks, matched Qwen general and coder models reliably raise Bloom cognitive demand but fail to lower it, so execution skill does not imply educational control.

  4. Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A 70-row table compiled from ophthalmology guidelines can label 3,000 synthetic dialogues, and fine-tuning a 9B model on them improves agreement with an author-defined reference from 61.7% to 74.1% and emergent recall...

  5. Deprecating Benchmarks: Criteria and Framework

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.

  6. Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances

    cs.CY 2025-07 accept novelty 6.0 of 10

    A Fair Equality of Chances-based framework decomposes GenAI unfairness into harms/benefits, morally arbitrary factors, and morally decisive factors to improve measurement validity.

  7. Establishing Best Practices for Building Rigorous Agentic Benchmarks

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Agentic benchmarks frequently mis-grade agents, and the new ABC checklist helps identify and correct such errors in ten popular benchmarks.

  8. Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Shortcut neuron patching suppresses benchmark-contamination shortcuts in LLMs and yields evaluation scores that strongly correlate with the external MixEval benchmark.

  9. Large Language Models in Code Co-generation for Safe Autonomous Vehicles

    cs.SE 2025-05 conditional novelty 6.0 of 10

    In a 24-setup benchmark using the esmini simulator, GPT-4 was the only model to produce a passing collision-avoidance controller, and it succeeded only once in 20 runs.

  10. Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Fine-tuning an LLM recommender on a slice of the benchmark inflates AUC/UAUC for in-domain leakage and degrades it for out-of-domain leakage, showing benchmark contamination can distort LLM-based recommendation evaluation.

  11. InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A submodular mutual information framework for selecting and training in-context learning exemplars improves average accuracy on nine benchmarks by about five points over the IDEAL baseline.

  12. ISACL: Internal State Analyzer for Copyrighted Training Data Leakage

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An MLP trained on LLM internal states predicts Rouge-L-defined literal copying leakage with high accuracy, but not paraphrase-level leakage.

  13. User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors introduce an entropy-based framework that uses user behavior prediction as a measure of LLM generalization, and find GPT-4o outperforms GPT-4o-mini and Llama-3.1 on movie and music recommendation tasks.

  14. A Conceptual Framework for AI Capability Evaluations

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.

Pith tools