REVIEW 4 cited by
Generative Table Pre-training Empowers Models for Tabular Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recently, the topic of table pre-training has attracted considerable research interest. However, how to employ table pre-training to boost the performance of tabular prediction remains an open challenge. In this paper, we propose TapTap, the first attempt that leverages table pre-training to empower models for tabular prediction. After pre-training on a large corpus of real-world tabular data, TapTap can generate high-quality synthetic tables to support various applications on tabular data, including privacy protection, low resource regime, missing value imputation, and imbalanced classification. Extensive experiments on 12 datasets demonstrate that TapTap outperforms a total of 16 baselines in different scenarios. Meanwhile, it can be easily combined with various backbone models, including LightGBM, Multilayer Perceptron (MLP) and Transformer. Moreover, with the aid of table pre-training, models trained using synthetic data generated by TapTap can even compete with models using the original dataset on half of the experimental datasets, marking a milestone in the development of synthetic tabular data generation. The codes are available at https://github.com/ZhangTP1996/TapTap.
Forward citations
Cited by 4 Pith papers
-
AIGT: AI Generative Table Based on Prompt
A prompt-enhanced language model with a column-partitioning algorithm generates synthetic tabular data that beats existing methods on 14 of 20 public datasets.
-
Zero-Shot Decision Tree Construction via Large Language Models
A prompt-based algorithm that constructs CART-style decision trees from feature descriptions alone, using LLM probability estimates instead of data.
-
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
A historical review that organizes multimodal explainability methods into four chronological eras and three explainability types, extending coverage to generative LLMs.
-
A Comprehensive Survey of Synthetic Tabular Data Generation
A structured survey that categorizes synthetic tabular data generation into traditional, diffusion, and LLM-based methods, with a comparative benchmark and a taxonomy of post-processing and evaluation.
Discussion (0). Continue with ORCID to comment.