Pith. sign in

REVIEW 18 cited by

On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15126 v1 pith:W5IJQHMB submitted 2024-06-14 cs.CL

classification cs.CL
keywords datagenerationsyntheticllms-drivenacademicadventaimsalleviate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Within the evolving landscape of deep learning, the dilemma of data quantity and quality has been a long-standing problem. The recent advent of Large Language Models (LLMs) offers a data-centric solution to alleviate the limitations of real-world data with synthetic data generation. However, current investigations into this field lack a unified framework and mostly stay on the surface. Therefore, this paper provides an organization of relevant studies based on a generic workflow of synthetic data generation. By doing so, we highlight the gaps within existing research and outline prospective avenues for future study. This work aims to shepherd the academic and industrial communities towards deeper, more methodical inquiries into the capabilities and applications of LLMs-driven synthetic data generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules

    cs.AI 2025-05 conditional novelty 6.5 of 10

    QuAda, a trainable attention adapter using under 2.8% extra parameters, gives instruction-tuned LLMs strong performance on five quotation-aware dialogue tasks.

  2. Evaluating Zero-Shot and One-Shot Adaptation of Small Language Models in Leader-Follower Interaction

    cs.HC 2026-02 conditional novelty 6.0 of 10

    Fine-tuned Qwen2.5-0.5B classifies leader-follower roles with 86.66% accuracy in single-turn interactions, but accuracy falls to chance in one-shot multi-turn interactions.

  3. Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).

  4. Ensembling LLM-Induced Decision Trees for Explainable and Robust Error Detection

    cs.CL 2025-12 conditional novelty 6.0 of 10

    LLM-induced hybrid decision trees (rules + trained graph checks) ensembled via EM detect erroneous table cells with an average 16.1-point F1 gain over the best baseline.

  5. StaAgent: An Agentic Framework for Testing Static Analyzers

    cs.SE 2025-07 conditional novelty 6.0 of 10

    An LLM-powered four-agent framework performs metamorphic testing on static analyzers and reports 64 faulty rule implementations across SpotBugs, SonarQube, ErrorProne, Infer, and PMD.

  6. Multimodal Mathematical Reasoning with Diverse Solving Perspective

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Training a multimodal language model on multiple diverse solution paths per problem, plus rewards for distinguishing correct from incorrect solutions, improves math benchmark accuracy and output diversity.

  7. MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MDBench is a synthetically generated, knowledge-guided benchmark for multi-document QA on which frontier LLMs achieve only about 60% exact match.

  8. RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions

    cs.CV 2025-06 conditional novelty 6.0 of 10

    RefEdit-Bench measures referring-expression image editing; the RefEdit model, trained on 20K synthetic triplets, reports state-of-the-art results over million-scale baselines.

  9. Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    UCerF scores LLM fairness by both correctness and confidence, and SynthBias provides 31,756 gender-occupation coreference samples for benchmark testing.

  10. Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Self-Error-Instruct clusters a model's math errors into types, synthesizes targeted practice data per type, and selects the best samples for fine-tuning, improving held-out math test accuracy.

  11. Can Structured Templates Facilitate LLMs in Tackling Harder Tasks? : An Exploration of Scaling Laws by Difficulty

    cs.AI 2025-08 reject novelty 5.0 of 10

    Training on easy synthetic math data lowers accuracy on hard benchmarks, and the proposed SST framework, which teaches explicit procedural chains, aims to reverse that drop.

  12. Separation Logic of Generic Resources via Sheafeology

    cs.LO 2025-08 unverdicted novelty 5.0 of 10

    Sheafeology uses sheaf categories to make first-order logic resource-aware, yielding separation logics for generic resources such as memory and random variables.

  13. Enterprise Large Language Model Evaluation Benchmark

    cs.AI 2025-06 reject novelty 5.0 of 10

    A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...

  14. A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations

    cs.CL 2025-07 conditional novelty 4.0 of 10

    PATTR adds a target-length penalty to the Type-Token Ratio, producing a lexical diversity score with tunable, reduced short-text bias for LLM synthetic data.

  15. Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead

    cs.SE 2025-06 conditional novelty 4.0 of 10

    A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.

  16. Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration

    eess.SP 2025-06 conditional novelty 4.0 of 10

    The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...

  17. Does Prompt Design Impact Quality of Data Imputation by LLMs?

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Group-wise CSV prompts with correlation-based column pruning reduce LLM imputation prompt size while roughly maintaining or slightly improving classifier-based imputation quality on two imbalanced datasets.

  18. Unlocking the Potential of Large Language Models in the Nuclear Industry with Synthetic Data

    cs.CL 2025-06 conditional novelty 2.0 of 10

    A pipeline converts CANDU textbook chapters into synthetic QA pairs using LLMs, embedding clustering, and similarity metrics, with no downstream validation.

Pith tools