Pith. sign in

REVIEW 6 cited by

ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.13048 v2 pith:URI4PAGD submitted 2020-12-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagenaturalimplicationsproofsmethodsproofproofwritertheory
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have been shown to emulate logical deduction over natural language theories (logical rules expressed in natural language), reliably assigning true/false labels to candidate implications. However, their ability to generate implications of a theory has not yet been demonstrated, and methods for reconstructing proofs of answers are imperfect. In this work we show that a generative model, called ProofWriter, can reliably generate both implications of a theory and the natural language proof(s) that support them. In particular, iterating a 1-step implication generator results in proofs that are highly reliable, and represent actual model decisions (rather than post-hoc rationalizations). On the RuleTaker dataset, the accuracy of ProofWriter's proofs exceed previous methods by +9% absolute, and in a way that generalizes to proof depths unseen in training and on out-of-domain problems. We also show that generative techniques can perform a type of abduction with high precision: Given a theory and an unprovable conclusion, identify a missing fact that allows the conclusion to be proved, along with a proof. These results significantly improve the viability of neural methods for systematically reasoning over natural language.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Long chain-of-thought and RL training on math problems improves general reasoning benchmarks, while short chain-of-thought math fine-tuning often degrades performance.

  2. Reasoning is about giving reasons

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A transformer can convert reasoning sentences into a task-specific logical form with 95-99% exact match, after which a symbolic solver answers deductive queries.

  3. Mitigating Spurious Correlations in LLMs via Causality-Aware Post-Training

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fine-tuning a 3B LLM on randomly symbolized reasoning questions reduces spurious-correlation failures and improves OOD accuracy on CLadder and PrOntoQA.

  4. LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A linear probe over layer-wise logit-lens probabilities ranks LLM confidence well enough to slightly beat voting or probability-based baselines in QA ensembles.

  5. Synergizing Logical Reasoning, Knowledge Management and Collaboration in Multi-Agent LLM System

    cs.MA 2025-07 conditional novelty 4.0 of 10

    SynergyMAS combines a graph database with a Clingo logic solver, corrective RAG, and Theory of Mind prompts in a hierarchical multi-agent team, demonstrated on a Smart Home Energy Management case study.

  6. Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models

    cs.AI 2025-06 reject novelty 3.0 of 10

    A 4B-parameter model is claimed to explain its own reasoning through inverse attention analysis, but the paper offers no consistent evidence or artifacts.

Pith tools