REVIEW 18 cited by
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Within the evolving landscape of deep learning, the dilemma of data quantity and quality has been a long-standing problem. The recent advent of Large Language Models (LLMs) offers a data-centric solution to alleviate the limitations of real-world data with synthetic data generation. However, current investigations into this field lack a unified framework and mostly stay on the surface. Therefore, this paper provides an organization of relevant studies based on a generic workflow of synthetic data generation. By doing so, we highlight the gaps within existing research and outline prospective avenues for future study. This work aims to shepherd the academic and industrial communities towards deeper, more methodical inquiries into the capabilities and applications of LLMs-driven synthetic data generation.
Forward citations
Cited by 18 Pith papers
-
Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules
QuAda, a trainable attention adapter using under 2.8% extra parameters, gives instruction-tuned LLMs strong performance on five quotation-aware dialogue tasks.
-
Evaluating Zero-Shot and One-Shot Adaptation of Small Language Models in Leader-Follower Interaction
Fine-tuned Qwen2.5-0.5B classifies leader-follower roles with 86.66% accuracy in single-turn interactions, but accuracy falls to chance in one-shot multi-turn interactions.
-
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).
-
Ensembling LLM-Induced Decision Trees for Explainable and Robust Error Detection
LLM-induced hybrid decision trees (rules + trained graph checks) ensembled via EM detect erroneous table cells with an average 16.1-point F1 gain over the best baseline.
-
StaAgent: An Agentic Framework for Testing Static Analyzers
An LLM-powered four-agent framework performs metamorphic testing on static analyzers and reports 64 faulty rule implementations across SpotBugs, SonarQube, ErrorProne, Infer, and PMD.
-
Multimodal Mathematical Reasoning with Diverse Solving Perspective
Training a multimodal language model on multiple diverse solution paths per problem, plus rewards for distinguishing correct from incorrect solutions, improves math benchmark accuracy and output diversity.
-
MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance
MDBench is a synthetically generated, knowledge-guided benchmark for multi-document QA on which frontier LLMs achieve only about 60% exact match.
-
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
RefEdit-Bench measures referring-expression image editing; the RefEdit model, trained on 20K synthetic triplets, reports state-of-the-art results over million-scale baselines.
-
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
UCerF scores LLM fairness by both correctness and confidence, and SynthBias provides 31,756 gender-occupation coreference samples for benchmark testing.
-
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
Self-Error-Instruct clusters a model's math errors into types, synthesizes targeted practice data per type, and selects the best samples for fine-tuning, improving held-out math test accuracy.
-
Can Structured Templates Facilitate LLMs in Tackling Harder Tasks? : An Exploration of Scaling Laws by Difficulty
Training on easy synthetic math data lowers accuracy on hard benchmarks, and the proposed SST framework, which teaches explicit procedural chains, aims to reverse that drop.
-
Separation Logic of Generic Resources via Sheafeology
Sheafeology uses sheaf categories to make first-order logic resource-aware, yielding separation logics for generic resources such as memory and random variables.
-
Enterprise Large Language Model Evaluation Benchmark
A 14-task enterprise LLM benchmark built mostly from GPT-4o-generated labels and scored by GPT-4o-as-judge shows open-source models closing the reasoning gap, but the dataset is not public and the evaluation is partly...
-
A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
PATTR adds a target-length penalty to the Type-Token Ratio, producing a lexical diversity score with tunable, reduced short-text bias for LLM synthetic data.
-
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.
-
Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...
-
Does Prompt Design Impact Quality of Data Imputation by LLMs?
Group-wise CSV prompts with correlation-based column pruning reduce LLM imputation prompt size while roughly maintaining or slightly improving classifier-based imputation quality on two imbalanced datasets.
-
Unlocking the Potential of Large Language Models in the Nuclear Industry with Synthetic Data
A pipeline converts CANDU textbook chapters into synthetic QA pairs using LLMs, embedding clustering, and similarity metrics, with no downstream validation.
Discussion (0). Sign in to comment.