REVIEW 5 cited by
LLMs May Perform MCQA by Selecting the Least Incorrect Option
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In the field of NLP, Large Language Models (LLMs) have markedly enhanced performance across a variety of tasks. However, the comprehensive evaluation of LLMs remains an inevitable challenge for the community. Recently, the adoption of Multiple Choice Question Answering (MCQA) as a benchmark for assessing LLMs has gained considerable traction. However, concerns regarding the robustness of this evaluative method persist. Building upon previous discussions on the issue of \textit{variability}, we reveal an additional dimension of concern: LLMs may perform MCQA by selecting the least incorrect option rather than distinctly correct. This observation suggests that LLMs might regard multiple options as correct, which could undermine the reliability of MCQA as a metric for evaluating LLMs. To address this challenge, we introduce an enhanced dataset augmentation method for MCQA, termed MCQA+, to provide a more accurate reflection of the model performance, thereby highlighting the necessity for more sophisticated evaluation mechanisms in the assessment of LLM capabilities.
Forward citations
Cited by 5 Pith papers
-
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
Map-reduce scaffolding degrades measured safety mainly by stripping multiple-choice options (40–89% of the loss is format conversion); scaffold architecture explains only 0.4% of variance and composite safety scores h...
-
The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.
-
Impact of Comments on LLM Comprehension of Legacy Code
Increased comment prevalence improved LLM quiz accuracy on legacy assembler code, while completely inaccurate comments degraded accuracy by roughly 12 percentage points.
-
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
ZeroSumEval ranks 13 LLMs through over 7,000 head-to-head games and finds models fall short at creative generation and jailbreaking.
-
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic
A 1.6B Arabic language model trained with synthetic multiple-choice instruction data outperforms 7B-13B models on several Arabic multiple-choice benchmarks.
Discussion (0). Continue with ORCID to comment.