Pith. sign in

REVIEW 6 cited by

Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02990 v4 pith:5SFVAK6B submitted 2024-03-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords dataaugmentationllmschallengeslanguagelearninghighlightslarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the rapidly evolving field of large language models (LLMs), data augmentation (DA) has emerged as a pivotal technique for enhancing model performance by diversifying training examples without the need for additional data collection. This survey explores the transformative impact of LLMs on DA, particularly addressing the unique challenges and opportunities they present in the context of natural language processing (NLP) and beyond. From both data and learning perspectives, we examine various strategies that utilize LLMs for data augmentation, including a novel exploration of learning paradigms where LLM-generated data is used for diverse forms of further training. Additionally, this paper highlights the primary open challenges faced in this domain, ranging from controllable data augmentation to multi-modal data augmentation. This survey highlights a paradigm shift introduced by LLMs in DA, and aims to serve as a comprehensive guide for researchers and practitioners.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism Detection

    cs.CL 2025-06 conditional novelty 6.0 of 10

    On the EDOS benchmark, definition-based augmentation and context expansion with a Mistral-7B tie-breaker reach macro F1 0.8819 (binary) and 0.6018 (fine-grained).

  2. LLMSynthor: Macro-Aligned Micro-Records Synthesis with Large Language Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMSynthor iteratively prompts an LLM to propose corrective batches of micro-records, aligning synthetic data with target macro-statistics while preserving realistic joint dependencies.

  3. Measuring Diversity in Synthetic Datasets

    cs.CL 2025-02 conditional novelty 6.0 of 10

    DCScore measures dataset diversity as the sum of self-classification probabilities under a softmax similarity matrix, and the paper shows it tracks generation temperature, human judgment, and LLM rankings.

  4. Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing

    cs.LG 2025-09 conditional novelty 4.0 of 10

    LLM tutoring modestly accelerates RL convergence on average, with advice reuse saving wall-clock time but reducing stability.

  5. Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A comprehensive survey of cross-lingual aspect-based sentiment analysis that catalogs tasks, datasets, modeling paradigms, and cross-lingual transfer techniques, and identifies research gaps.

  6. Large Language Models in the Data Science Lifecycle: A Systematic Mapping Study

    cs.CY 2025-08 conditional novelty 4.0 of 10

    A systematic mapping study classifying 66 papers on LLM use across five data science lifecycle stages, finding data analysis most studied and deployment almost ignored.

Pith tools