Pith. sign in

REVIEW 1 cited by

DAGA: Data Augmentation with a Generation Approach for Low-resource Tagging Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.01549 v1 pith:RDYOSPRI submitted 2020-11-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords datamethodaugmentationsettingstaggingtasksgivenlow-resource
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Data augmentation techniques have been widely used to improve machine learning performance as they enhance the generalization capability of models. In this work, to generate high quality synthetic data for low-resource tagging tasks, we propose a novel augmentation method with language models trained on the linearized labeled sentences. Our method is applicable to both supervised and semi-supervised settings. For the supervised settings, we conduct extensive experiments on named entity recognition (NER), part of speech (POS) tagging and end-to-end target based sentiment analysis (E2E-TBSA) tasks. For the semi-supervised settings, we evaluate our method on the NER task under the conditions of given unlabeled data only and unlabeled data plus a knowledge base. The results show that our method can consistently outperform the baselines, particularly when the given gold training data are less.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Less is More: Adaptive Coverage for Synthetic Training Data

    cs.LG 2025-04 conditional novelty 5.0 of 10

    A max-coverage graph algorithm with an adaptive similarity threshold selects 10-30% of synthetic data that trains classifiers as well as or better than the full dataset on sentiment, relation extraction, and NER tasks.

Pith tools