Pith. sign in

REVIEW 2 cited by

A Survey on Self-Supervised Learning for Non-Sequential Tabular Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01204 v4 pith:R5OGWDMH submitted 2024-02-02 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningtabulardatassl4ns-tdchallengesdatasetsdomainexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised learning (SSL) has been incorporated into many state-of-the-art models in various domains, where SSL defines pretext tasks based on unlabeled datasets to learn contextualized and robust representations. Recently, SSL has become a new trend in exploring the representation learning capability in the realm of tabular data, which is more challenging due to not having explicit relations for learning descriptive representations. This survey aims to systematically review and summarize the recent progress and challenges of SSL for non-sequential tabular data (SSL4NS-TD). We first present a formal definition of NS-TD and clarify its correlation to related studies. Then, these approaches are categorized into three groups - predictive learning, contrastive learning, and hybrid learning, with their motivations and strengths of representative methods in each direction. Moreover, application issues of SSL4NS-TD are presented, including automatic data engineering, cross-table transferability, and domain knowledge integration. In addition, we elaborate on existing benchmarks and datasets for NS-TD applications to analyze the performance of existing tabular models. Finally, we discuss the challenges of SSL4NS-TD and provide potential directions for future research. We expect our work to be useful in terms of encouraging more research on lowering the barrier to entry SSL for the tabular domain, and of improving the foundations for implicit tabular data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ORIGAMI: A generative transformer architecture for predictions from semi-structured data

    cs.LG 2024-12 conditional novelty 6.0 of 10

    ORIGAMI is a generative transformer with structure-preserving tokenization, key/value position encoding, and grammar-constrained decoding that matches or beats baselines on tabular, multi-label, and code-classification tasks.

  2. APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    APAR pre-trains a tabular transformer on arithmetic combinations of target labels and fine-tunes it with adaptive feature masking, beating GBDT and neural baselines on 10 regression datasets.

Pith tools