Pith. sign in

REVIEW 2 cited by

Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14279 v4 pith:PKRDOS4R submitted 2023-05-23 cs.CL

classification cs.CL
keywords consistencymodelmulti-stepreasoningself-consistencytaskshypotheticalimportant
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have achieved widespread success on a variety of in-context few-shot tasks, but this success is typically evaluated via correctness rather than consistency. We argue that self-consistency is an important criteria for valid multi-step reasoning in tasks where the solution is composed of the answers to multiple sub-steps. We propose two types of self-consistency that are particularly important for multi-step reasoning -- hypothetical consistency (a model's ability to predict what its output would be in a hypothetical other context) and compositional consistency (consistency of a model's final outputs when intermediate sub-steps are replaced with the model's outputs for those steps). We demonstrate that multiple variants of the GPT-3/-4 models exhibit poor consistency rates across both types of consistency on a variety of tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Large language models are less logically consistent when hypotheses are decomposed into atomic sub-problems, and a new inferential-consistency metric quantifies how consistently models handle the same fact in differen...

  2. Does It Make Sense to Speak of Introspection in Large Language Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors argue that an untrained large language model inferring its own sampling temperature from the style of its own output qualifies as a minimal, consciousness-free form of introspection.

Pith tools