Pith. sign in

REVIEW 5 cited by

Abductive Commonsense Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.05739 v2 pith:SDLSWXUS submitted 2019-08-15 cs.CL

classification cs.CL
keywords abductivelanguagereasoningexplanationnaturaltaskbeenbest
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Abductive reasoning is inference to the most plausible explanation. For example, if Jenny finds her house in a mess when she returns from work, and remembers that she left a window open, she can hypothesize that a thief broke into her house and caused the mess, as the most plausible explanation. While abduction has long been considered to be at the core of how people interpret and read between the lines in natural language (Hobbs et al., 1988), there has been relatively little research in support of abductive natural language inference and generation. We present the first study that investigates the viability of language-based abductive reasoning. We introduce a challenge dataset, ART, that consists of over 20k commonsense narrative contexts and 200k explanations. Based on this dataset, we conceptualize two new tasks -- (i) Abductive NLI: a multiple-choice question answering task for choosing the more likely explanation, and (ii) Abductive NLG: a conditional generation task for explaining given observations in natural language. On Abductive NLI, the best model achieves 68.9% accuracy, well below human performance of 91.4%. On Abductive NLG, the current best language generators struggle even more, as they lack reasoning capabilities that are trivial for humans. Our analysis leads to new insights into the types of reasoning that deep pre-trained language models fail to perform--despite their strong performance on the related but more narrowly defined task of entailment NLI--pointing to interesting avenues for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

    cs.CL 2026-07 conditional novelty 6.0 of 10

    An NLI-hypergraph audit with deterministic AND–OR search labels LLM reasoning segments as supported, unsupported, or orphaned, improving balanced F1 over LLM-as-judge on a new 40-case clinical benchmark.

  2. Entailed Between the Lines: Incorporating Implication into NLI

    cs.CL 2025-01 conditional novelty 6.0 of 10

    The paper formalizes implied entailment as a four-way NLI label, builds the INLI dataset from existing implicature sources, and shows fine-tuned models can label implied vs explicit entailment and generalize across co...

  3. Enhancing Transformers for Generalizable First-Order Logical Entailment

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Transformers with relative positional encoding beat KGQA baselines, and adding logic-aware attention (TEGA) improves out-of-distribution performance on a new 55-type benchmark.

  4. Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing

    cs.CV 2025-01 conditional novelty 5.0 of 10

    G2 generates a location-free scene graph from an image and feeds it, with confidence-based token weighting, into an LLM to produce visual commonsense answers and explanations.

  5. Advancing Reasoning in Large Language Models: Promising Methods and Approaches

    cs.CL 2025-02 conditional

    A survey that categorizes existing LLM reasoning techniques into prompting, architectural, and learning-based approaches, without contributing new results.

Pith tools