Pith. sign in

REVIEW 3 cited by

LLM-Generated Natural Language Meets Scaling Laws: New Explorations and Data Augmentation Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00322 v1 pith:7F2QMH3Z submitted 2024-06-29 cs.CL

classification cs.CL
keywords datalanguagenaturalaugmentationlawsllmnlscalingaugmented
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the ascent of large language models (LLM), natural language processing has witnessed enhancements, such as LLM-based data augmentation. Nonetheless, prior research harbors two primary concerns: firstly, a lack of contemplation regarding whether the natural language generated by LLM (LLMNL) truly aligns with human natural language (HNL), a critical foundational question; secondly, an oversight that augmented data is randomly generated by LLM, implying that not all data may possess equal training value, that could impede the performance of classifiers. To address these challenges, we introduce the scaling laws to intrinsically calculate LLMNL and HNL. Through extensive experiments, we reveal slight deviations (approximately 0.2 Mandelbrot exponent) from Mandelbrot's law in LLMNL, underscore a complexity advantage in HNL, and supplement an interpretive discussion on language style. This establishes a solid foundation for LLM's expansion. Further, we introduce a novel data augmentation method for few-shot text classification, termed ZGPTDA, which leverages fuzzy computing mechanisms driven by the conformity to scaling laws to make decisions about GPT-4 augmented data. Extensive experiments, conducted in real-world scenarios, confirms the effectiveness (improving F1 of Bert and RoBerta by 7-10%) and competitiveness (surpassing recent AugGPT and GENCO methods by about 2% accuracy on DeBerta) of ZGPTDA. In addition, we reveal some interesting insights, e.g., Hilberg's law and Taylor's law can impart more benefits to text classification, etc.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GOBench measures how well multimodal AI models generate and understand geometric optics, finding that even top models make frequent physical errors.

  2. Scaling-up Perceptual Video Quality Assessment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new pipeline plus datasets (OmniVQA-Chat-400K, OmniVQA-MOS-20K, OmniVQA-FG-Benchmark) yield LMMs with state-of-the-art video quality understanding and rating.

  3. ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A fine-tuned LLaMA-2-7B model trained with selected LLM-generated data improves extraction of chemical synthesis actions from experimental text.

Pith tools