REVIEW 6 cited by
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the realm of Computational Social Science (CSS), practitioners often navigate complex, low-resource domains and face the costly and time-intensive challenges of acquiring and annotating data. We aim to establish a set of guidelines to address such challenges, comparing the use of human-labeled data with synthetically generated data from GPT-4 and Llama-2 in ten distinct CSS classification tasks of varying complexity. Additionally, we examine the impact of training data sizes on performance. Our findings reveal that models trained on human-labeled data consistently exhibit superior or comparable performance compared to their synthetically augmented counterparts. Nevertheless, synthetic augmentation proves beneficial, particularly in improving performance on rare classes within multi-class tasks. Furthermore, we leverage GPT-4 and Llama-2 for zero-shot classification and find that, while they generally display strong performance, they often fall short when compared to specialized classifiers trained on moderately sized training sets.
Forward citations
Cited by 6 Pith papers
-
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
A joint diffusion framework trains a Stable Diffusion generator and a semantic occupancy perception model together, so each task improves the other, producing text-conditional RGB-occupancy pairs.
-
Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
Backtranslation and paraphrasing produce competitive or better classification gains than zero-shot and few-shot generation when augmenting a low-resource emotion dataset.
-
Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration
RAC ranks label descriptions by semantic similarity, then runs iterative binary LLM checks one label at a time, improving zero-shot labelling accuracy and enabling a precision-coverage trade-off.
-
Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method
UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...
-
Automated Extraction of Acronym-Expansion Pairs from Scientific Papers
A regex-plus-GPT-4 pipeline extracts acronym-expansion pairs from scientific papers more fully than either component alone, though precision is not measured.
-
Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems
A label taxonomy and answer-first generation strategies help RAG developers build evaluation datasets whose question mix matches real usage.
Discussion (0). Continue with ORCID to comment.