Pith. sign in

REVIEW 1 cited by

Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.11472 v3 pith:JZNJ5SW4 submitted 2022-05-23 cs.CL

classification cs.CL
keywords argumentminingdatasetstaskcomponentsdatasetdifficulteffect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The task of Argument Mining, that is extracting and classifying argument components for a specific topic from large document sources, is an inherently difficult task for machine learning models and humans alike, as large Argument Mining datasets are rare and recognition of argument components requires expert knowledge. The task becomes even more difficult if it also involves stance detection of retrieved arguments. In this work, we investigate the effect of Argument Mining dataset composition in few- and zero-shot settings. Our findings show that, while fine-tuning is mandatory to achieve acceptable model performance, using carefully composed training samples and reducing the training sample size by up to almost 90% can still yield 95% of the maximum performance. This gain is consistent across three Argument Mining tasks on three different datasets. We also publish a new dataset for future benchmarking.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Rationalization in the Wild: A Large Scale Out-of-Distribution Evaluation on NLI-related tasks

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Fine-tuning self-rationalization models on few examples transfers to 19 OOD NLI-related datasets nearly as well as full-data fine-tuning, and the Acceptability score is the best reference-free explanation metric tested.

Pith tools