REVIEW 6 cited by
Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the rapidly evolving field of large language models (LLMs), data augmentation (DA) has emerged as a pivotal technique for enhancing model performance by diversifying training examples without the need for additional data collection. This survey explores the transformative impact of LLMs on DA, particularly addressing the unique challenges and opportunities they present in the context of natural language processing (NLP) and beyond. From both data and learning perspectives, we examine various strategies that utilize LLMs for data augmentation, including a novel exploration of learning paradigms where LLM-generated data is used for diverse forms of further training. Additionally, this paper highlights the primary open challenges faced in this domain, ranging from controllable data augmentation to multi-modal data augmentation. This survey highlights a paradigm shift introduced by LLMs in DA, and aims to serve as a comprehensive guide for researchers and practitioners.
Forward citations
Cited by 6 Pith papers
-
Explaining Matters: Leveraging Definitions and Semantic Expansion for Sexism Detection
On the EDOS benchmark, definition-based augmentation and context expansion with a Mistral-7B tie-breaker reach macro F1 0.8819 (binary) and 0.6018 (fine-grained).
-
LLMSynthor: Macro-Aligned Micro-Records Synthesis with Large Language Models
LLMSynthor iteratively prompts an LLM to propose corrective batches of micro-records, aligning synthetic data with target macro-statistics while preserving realistic joint dependencies.
-
Measuring Diversity in Synthetic Datasets
DCScore measures dataset diversity as the sum of self-classification probabilities under a softmax similarity matrix, and the paper shows it tracks generation temperature, human judgment, and LLM rankings.
-
Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing
LLM tutoring modestly accelerates RL convergence on average, with advice reuse saving wall-clock time but reducing stability.
-
Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
A comprehensive survey of cross-lingual aspect-based sentiment analysis that catalogs tasks, datasets, modeling paradigms, and cross-lingual transfer techniques, and identifies research gaps.
-
Large Language Models in the Data Science Lifecycle: A Systematic Mapping Study
A systematic mapping study classifying 66 papers on LLM use across five data science lifecycle stages, finding data analysis most studied and deployment almost ignored.
Discussion (0). Continue with ORCID to comment.