Pith. sign in

REVIEW 5 cited by

LLMs May Perform MCQA by Selecting the Least Incorrect Option

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01349 v3 pith:3VAR6D3M submitted 2024-02-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsmcqachallengecorrectenhancedevaluationhoweverincorrect
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In the field of NLP, Large Language Models (LLMs) have markedly enhanced performance across a variety of tasks. However, the comprehensive evaluation of LLMs remains an inevitable challenge for the community. Recently, the adoption of Multiple Choice Question Answering (MCQA) as a benchmark for assessing LLMs has gained considerable traction. However, concerns regarding the robustness of this evaluative method persist. Building upon previous discussions on the issue of \textit{variability}, we reveal an additional dimension of concern: LLMs may perform MCQA by selecting the least incorrect option rather than distinctly correct. This observation suggests that LLMs might regard multiple options as correct, which could undermine the reliability of MCQA as a metric for evaluating LLMs. To address this challenge, we introduce an enhanced dataset augmentation method for MCQA, termed MCQA+, to provide a more accurate reflection of the model performance, thereby highlighting the necessity for more sophisticated evaluation mechanisms in the assessment of LLM capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

    cs.SE 2026-03 conditional novelty 7.0 of 10

    Map-reduce scaffolding degrades measured safety mainly by stripping multiple-choice options (40–89% of the loss is format conversion); scaffold architecture explains only 0.4% of variance and composite safety scores h...

  2. The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.

  3. Impact of Comments on LLM Comprehension of Legacy Code

    cs.SE 2025-04 conditional novelty 5.0 of 10

    Increased comment prevalence improved LLM quiz accuracy on legacy assembler code, while completely inaccurate comments degraded accuracy by roughly 12 percentage points.

  4. ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition

    cs.AI 2025-04 conditional novelty 5.0 of 10

    ZeroSumEval ranks 13 LLMs through over 7,000 head-to-head games and finds models fall short at creative generation and jailbreaking.

  5. Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A 1.6B Arabic language model trained with synthetic multiple-choice instruction data outperforms 7B-13B models on several Arabic multiple-choice benchmarks.

Pith tools