Pith. sign in

REVIEW 5 cited by

Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03439 v3 pith:C5CZKSBZ submitted 2023-04-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoninggpt-4logicaldatasetschatgptbenchmarksperformancelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as "advanced" at reasoning tasks, we are eager to learn the GPT-4 performance on various logical reasoning tasks. This report analyses multiple logical reasoning datasets, with popular benchmarks like LogiQA and ReClor, and newly-released datasets like AR-LSAT. We test the multi-choice reading comprehension and natural language inference tasks with benchmarks requiring logical reasoning. We further construct a logical reasoning out-of-distribution dataset to investigate the robustness of ChatGPT and GPT-4. We also make a performance comparison between ChatGPT and GPT-4. Experiment results show that ChatGPT performs significantly better than the RoBERTa fine-tuning method on most logical reasoning benchmarks. With early access to the GPT-4 API we are able to conduct intense experiments on the GPT-4 model. The results show GPT-4 yields even higher performance on most logical reasoning datasets. Among benchmarks, ChatGPT and GPT-4 do relatively well on well-known datasets like LogiQA and ReClor. However, the performance drops significantly when handling newly released and out-of-distribution datasets. Logical reasoning remains challenging for ChatGPT and GPT-4, especially on out-of-distribution and natural language inference datasets. We release the prompt-style logical reasoning datasets as a benchmark suite and name it LogiEval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.

  2. Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DZEN, a parallel Dzongkha-English benchmark of 5,161 school science exam questions, shows large LLM accuracy gaps in Dzongkha; adding English translations narrows the gap.

  3. NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Large language models are less logically consistent when hypotheses are decomposed into atomic sub-problems, and a new inferential-consistency metric quantifies how consistently models handle the same fact in differen...

  4. Self-Rationalization in the Wild: A Large Scale Out-of-Distribution Evaluation on NLI-related tasks

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Fine-tuning self-rationalization models on few examples transfers to 19 OOD NLI-related datasets nearly as well as full-data fine-tuning, and the Acceptability score is the best reference-free explanation metric tested.

  5. S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency

    cs.CL 2025-02 conditional novelty 5.0 of 10

    S2-MAD's decision mechanism filters redundant viewpoints and conditionally skips participation, cutting token costs by up to 94.5% versus standard multi-agent debate while keeping accuracy within about 2 points in the...

Pith tools