Pith. sign in

REVIEW 4 cited by

Faithful Reasoning Using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.14271 v1 pith:SKZPINNY submitted 2022-08-30 cs.AI cs.CL

classification cs.AIcs.CL
keywords reasoningmulti-stepdemonstratefaithfullanguagelargelogicalmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although contemporary large language models (LMs) demonstrate impressive question-answering capabilities, their answers are typically the product of a single call to the model. This entails an unwelcome degree of opacity and compromises performance, especially on problems that are inherently multi-step. To address these limitations, we show how LMs can be made to perform faithful multi-step reasoning via a process whose causal structure mirrors the underlying logical structure of the problem. Our approach works by chaining together reasoning steps, where each step results from calls to two fine-tuned LMs, one for selection and one for inference, to produce a valid reasoning trace. Our method carries out a beam search through the space of reasoning traces to improve reasoning quality. We demonstrate the effectiveness of our model on multi-step logical deduction and scientific question-answering, showing that it outperforms baselines on final answer accuracy, and generates humanly interpretable reasoning traces whose validity can be checked by the user.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Hallucinated captions systematically improve VLM accuracy on vision-language tasks across nine models and nine datasets, with gains linked to broadened semantic coverage and modulated reasoning entropy.

  2. Finetuning Lightweight LLMs for Control Flow Graph Generation

    cs.SE 2026-07 conditional novelty 5.5 of 10

    Fine-tuned ~3–7B LLMs generate unified digraph CFGs from incomplete/erroneous code and show partial cross-language transfer to held-out JavaScript.

  3. Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

    cs.SE 2025-06 accept novelty 5.0 of 10

    A new survey organizes LLM interpretation methods by workflow stage and connects them to safety enhancement strategies and tools, covering around 70 works.

  4. Two-way Evidence self-Alignment based Dual-Gated Reasoning Enhancement

    cs.CL 2025-05 conditional novelty 4.0 of 10

    ESA-DGR combines two-way evidence self-alignment with dual-gated knowledge fusion and GRPO training to improve multi-hop question answering on HotpotQA, 2WikiMultiHopQA, and MuSiQue.

Pith tools