Pith. sign in

REVIEW 2 cited by

e-SNLI-VE: Corrected Visual-Textual Entailment with Natural Language Explanations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.03744 v3 pith:VWOAJQQV submitted 2020-04-07 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords corpusexplanationssnli-vecorrectede-snli-veentailmentlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recently proposed SNLI-VE corpus for recognising visual-textual entailment is a large, real-world dataset for fine-grained multimodal reasoning. However, the automatic way in which SNLI-VE has been assembled (via combining parts of two related datasets) gives rise to a large number of errors in the labels of this corpus. In this paper, we first present a data collection effort to correct the class with the highest error rate in SNLI-VE. Secondly, we re-evaluate an existing model on the corrected corpus, which we call SNLI-VE-2.0, and provide a quantitative comparison with its performance on the non-corrected corpus. Thirdly, we introduce e-SNLI-VE, which appends human-written natural language explanations to SNLI-VE-2.0. Finally, we train models that learn from these explanations at training time, and output such explanations at testing time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Are Research Hypotheses?

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A position paper documenting inconsistent and often implicit definitions of 'hypothesis' across NLP hypothesis mining tasks and calling for standardization.

  2. Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Llama 3.2 Vision reaches 83.3% on e-SNLI-VE after fine-tuning, but high explanation scores persist with black images, showing VE accuracy and BERTScore are weak evidence of visual grounding.

Pith tools