Pith. sign in

REVIEW 4 cited by

Quick and (not so) Dirty: Unsupervised Selection of Justification Sentences for Multi-hop Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.07176 v2 pith:F5WWEG6J submitted 2019-11-17 cs.CL

classification cs.CL
keywords sentencesselectionjustificationselectedunsupervisedmethodmulti-hopmultirc
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose an unsupervised strategy for the selection of justification sentences for multi-hop question answering (QA) that (a) maximizes the relevance of the selected sentences, (b) minimizes the overlap between the selected facts, and (c) maximizes the coverage of both question and answer. This unsupervised sentence selection method can be coupled with any supervised QA approach. We show that the sentences selected by our method improve the performance of a state-of-the-art supervised QA model on two multi-hop QA datasets: AI2's Reasoning Challenge (ARC) and Multi-Sentence Reading Comprehension (MultiRC). We obtain new state-of-the-art performance on both datasets among approaches that do not use external resources for training the QA system: 56.82% F1 on ARC (41.24% on Challenge and 64.49% on Easy) and 26.1% EM0 on MultiRC. Our justification sentences have higher quality than the justifications selected by a strong information retrieval baseline, e.g., by 5.4% F1 in MultiRC. We also show that our unsupervised selection of justification sentences is more stable across domains than a state-of-the-art supervised sentence selection method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Delta Learning

    cs.LG 2026-01 unverdicted novelty 7.0 of 10

    Replacing additive residual connections with a gated rank-1 delta update that interpolates identity, projection, and reflection slightly improves language modeling and downstream averages in reported 124M/353M runs.

  2. JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    JustLogic, a synthetic benchmark with high linguistic and argument complexity, shows most LLMs underperform the average human in pure deductive reasoning.

  3. TruncFormer: Private LLM Inference Using Only Truncations

    cs.CR 2024-12 conditional novelty 6.0 of 10

    TruncFormer statically places truncations in private LLM inference so all nonlinear operations reduce to adds, multiplies, and truncations, cutting estimated truncation latency by up to about 1.92x versus PUMA without...

  4. InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion

    cs.CL 2025-01 reject novelty 5.0 of 10

    InfiFusion fuses multiple large language models into one pivot model using enhanced universal logit distillation, and reports that the fused model outperforms all source models on 11 benchmarks with a fraction of the ...

Pith tools