Pith. sign in

REVIEW 1 cited by

SampleLLM: Optimizing Tabular Data Synthesis in Recommendations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.16125 v3 pith:CDQGKENJ submitted 2025-01-27 cs.IR

classification cs.IR
keywords datadistributiontabularfeaturesamplellmsynthesislearningrecommendation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tabular data synthesis is crucial in machine learning, yet existing general methods-primarily based on statistical or deep learning models-are highly data-dependent and often fall short in recommender systems. This limitation arises from their difficulty in capturing complex distributions and understanding feature relationships from sparse and limited data, along with their inability to grasp semantic feature relations. Recently, Large Language Models (LLMs) have shown potential in generating synthetic data samples through few-shot learning and semantic understanding. However, they often suffer from inconsistent distribution and lack of diversity due to their inherent distribution disparity with the target dataset. To address these challenges and enhance tabular data synthesis for recommendation tasks, we propose a novel two-stage framework named SampleLLM to improve the quality of LLM-based tabular data synthesis for recommendations by ensuring better distribution alignment. In the first stage, SampleLLM employs LLMs with Chain-of-Thought prompts and diverse exemplars to generate data that closely aligns with the target dataset distribution, even when input samples are limited. The second stage uses an advanced feature attribution-based importance sampling method to refine feature relationships within the synthesized data, reducing any distribution biases introduced by the LLM. Experimental results on three recommendation datasets, two general datasets, and online deployment illustrate that SampleLLM significantly surpasses existing methods for recommendation tasks and holds promise for a broader range of tabular data scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Few-shot LLM Synthetic Data with Distribution Matching

    cs.CL 2025-02 conditional novelty 5.0 of 10

    SynAlign generates LLM synthetic text from diversity-guided demonstrations, then reweights it by MMD-based distribution matching, improving downstream classification accuracy.

Reference graph

Works this paper leans on

16 extracted references · 4 linked inside Pith · cited by 1 Pith paper

  1. [1]

    AI@Meta. 2024. Llama 3 Model Card. (2024)

  2. [2]

    Ms Aayushi Bansal, Dr Rewa Sharma, and Dr Mamta Kathuria. 2022. A systematic review on data scarcity problem in deep learning: solution and applications.ACM Computing Surveys (Csur) (2022), 1–29

  3. [3]

    Vadim Borisov, Tobias Leemann, Kathrin Seßler, Johannes Haug, Martin Pawel- czyk, and Gjergji Kasneci. 2022. Deep neural networks and tabular data: A survey. IEEE transactions on neural networks and learning systems (2022)

  4. [4]

    Vadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk, and Gjergji Kasneci. 2022. Language models are realistic tabular data generators. arXiv preprint arXiv:2210.06280 (2022)

  5. [5]

    Stavroula Bourou, Andreas El Saer, Terpsichori-Helen Velivassaki, Artemis Voulkidis, and Theodore Zahariadis. 2021. A review of tabular data synthesis using GANs on an IDS dataset. Information (2021), 375

  6. [6]

    Samuel Cahyawijaya, Holy Lovenia, and Pascale Fung. 2024. LLMs Are Few-Shot In-Context Low-Resource Language Learners. arXiv preprint arXiv:2403.16512 (2024)

  7. [7]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al . 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology (2024), 1–45

  8. [8]

    Subhajit Chatterjee, Debapriya Hazra, Yung-Cheol Byun, and Yong-Woon Kim

Show all 16 references
  1. [9]

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey on in-context learning. arXiv preprint arXiv:2301.00234 (2022)

  2. [10]

    Víctor Elvira and Luca Martino. 2021. Advances in importance sampling. arXiv preprint arXiv:2102.05407 (2021)

  3. [11]

    Gabriel Erion, Joseph D Janizek, Pascal Sturmfels, Scott M Lundberg, and Su-In Lee. [n. d.]. Learning explainable models using attribution priors. ([n. d.])

  4. [12]

    Sheikh Amir Fayaz, Majid Zaman, Sameer Kaul, and Muheet Ahmed Butt. 2022. Is deep learning on tabular data enough? An assessment. International Journal of Advanced Computer Science and Applications (2022)

  5. [13]

    Alvaro Figueira and Bruno Vaz. 2022. Survey on synthetic data generation, evaluation methods and GANs. Mathematics (2022), 2733

  6. [14]

    Joao Fonseca and Fernando Bacao. 2023. Tabular and latent space synthetic data generation: a literature review. Journal of Big Data (2023), 115

  7. [15]

    Zichuan Fu, Xiangyang Li, Chuhan Wu, Yichao Wang, Kuicai Dong, Xiangyu Zhao, Mengchen Zhao, Huifeng Guo, and Ruiming Tang. 2023. A unified frame- work for multi-domain ctr prediction via large language models. ACM Transac- tions on Information Systems (2023)

  8. [2022]

    Mathematics (2022), 1541

    Enhancement of image classification using transfer learning and GAN- based synthetic data augmentation. Mathematics (2022), 1541

Pith tools