Pith. sign in

REVIEW 16 cited by

SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.00436 v3 pith:MDFLHIBE submitted 2023-08-01 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords llmserrorsreasoningselfcheckanswerstep-by-stepzero-shotable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent progress in large language models (LLMs), especially the invention of chain-of-thought prompting, has made it possible to automatically answer questions by stepwise reasoning. However, when faced with more complicated problems that require non-linear thinking, even the strongest LLMs make mistakes. To address this, we explore whether LLMs are able to recognize errors in their own step-by-step reasoning, without resorting to external resources. To this end, we propose SelfCheck, a general-purpose zero-shot verification schema for recognizing such errors. We then use the results of these checks to improve question-answering performance by conducting weighted voting on multiple solutions to the question. We test SelfCheck on three datasets (GSM8K, MathQA, and MATH) and find that it successfully recognizes errors and, in turn, increases final answer accuracies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

    cs.AI 2026-07 conditional novelty 6.0 of 10

    LGU models implication and incompatibility among LLM answers and reports consistent AUROC/AUARC gains over semantic entropy on QA benchmarks.

  3. Using LLMs for Explainable, Data-Driven Insight Generation from Time Series

    cs.AI 2026-06 conditional novelty 6.0 of 10

    A modular LLM pipeline extracts explanatory factors from analyst text, conditions evidence-based report generation, and evaluates readability, consistency, and persuasiveness, reportedly matching analyst reports on re...

  4. Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Reward-guided optimization of SAE latents, gated by reward and confidence, improves multi-benchmark reasoning and is post-hoc associated with more verification and course correction.

  5. MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection

    cs.CL 2025-07 conditional novelty 6.0 of 10

    MIND uses unlabeled similar memes, bidirectional AI insight derivation, and multi-agent debate to improve zero-shot harmful meme detection on HarM, FHM, and MAMI.

  6. Leveraging Large Language Models for Tacit Knowledge Discovery in Organizational Contexts

    cs.AI 2025-07 conditional novelty 6.0 of 10

    An LLM agent can reconstruct full descriptions of data tables by conversing with simulated employees, even when the only full-knowledge expert is never contacted.

  7. AdapThink: Adaptive Thinking Preferences for Reasoning Language Model

    cs.LG 2025-06 conditional novelty 6.0 of 10

    AdapThink is an RL post-training framework that adaptively reduces overthinking and underthinking in reasoning language models by rewarding confidence-appropriate reasoning depth and diverse training samples.

  8. Your Agent Can Defend Itself against Backdoor Attacks

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A two-level consistency defense detects backdoored LLM agents by matching thoughts to actions and reconstructed instructions to the user's instruction, reducing attack success rates on tested tasks.

  9. Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A gradient-free Monte Carlo tree search over JSON key-step plans produces few-shot demonstrations that let LLaMA3-8B and LLaMA3.2-3B outperform GPT-3.5 on most of seven BIG-Bench Hard tasks.

  10. Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Knowledge probing methods for LLMs are found to be highly inconsistent under minor prompt perturbations and across different probes.

  11. Hume: Introducing System-2 Thinking in Visual-Language-Action Model

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A dual-system vision-language-action model that improves robot control by ranking multiple sampled action chunks with a learned value function before fast execution.

  12. Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference

    cs.AI 2025-09 reject novelty 5.0 of 10

    A hierarchical causal attribution framework (performance causal inversion, Shapley values, CDC-MAS step discovery) localizes failure agents and steps in LLM multi-agent systems, reporting up to 36.2% step accuracy.

  13. I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A text-first, multi-round visual feedback framework reports state-of-the-art top-1 accuracy on WikiMEL, WikiDiverse, and RichMEL.

  14. InfoFlood: Jailbreaking Large Language Models with Information Overload

    cs.CR 2025-06 conditional novelty 5.0 of 10

    InfoFlood claims near-perfect jailbreak success on four frontier LLMs by rewriting harmful queries into verbose academic prose with fake citations, past-tense framing, and ethical disclaimers, without adversarial suffixes.

  15. Generating on Generated: An Approach Towards Self-Evolving Diffusion Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A recursive self-training loop that filters prompts, selects preferred images, and reweights out-of-distribution samples improves Stable Diffusion models over multiple rounds.

  16. Beyond ROUGE: N-Gram Subspace Features for LLM Hallucination Detection

    cs.CL 2025-09 reject novelty 4.0 of 10

    Singular values of label-grouped n-gram frequency tensors are used as MLP features for hallucination detection, with reported gains on HaluEval that rely on label-aware grouping.

Pith tools