REVIEW 2 cited by
Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have achieved widespread success on a variety of in-context few-shot tasks, but this success is typically evaluated via correctness rather than consistency. We argue that self-consistency is an important criteria for valid multi-step reasoning in tasks where the solution is composed of the answers to multiple sub-steps. We propose two types of self-consistency that are particularly important for multi-step reasoning -- hypothetical consistency (a model's ability to predict what its output would be in a hypothetical other context) and compositional consistency (consistency of a model's final outputs when intermediate sub-steps are replaced with the model's outputs for those steps). We demonstrate that multiple variants of the GPT-3/-4 models exhibit poor consistency rates across both types of consistency on a variety of tasks.
Forward citations
Cited by 2 Pith papers
-
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
Large language models are less logically consistent when hypotheses are decomposed into atomic sub-problems, and a new inferential-consistency metric quantifies how consistently models handle the same fact in differen...
-
Does It Make Sense to Speak of Introspection in Large Language Models?
The authors argue that an untrained large language model inferring its own sampling temperature from the style of its own output qualifies as a minimal, consciousness-free form of introspection.
Discussion (0). Continue with ORCID to comment.