REVIEW 7 cited by
SCARF: Self-Supervised Contrastive Learning using Random Feature Corruption
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised contrastive representation learning has proved incredibly successful in the vision and natural language domains, enabling state-of-the-art performance with orders of magnitude less labeled data. However, such methods are domain-specific and little has been done to leverage this technique on real-world tabular datasets. We propose SCARF, a simple, widely-applicable technique for contrastive learning, where views are formed by corrupting a random subset of features. When applied to pre-train deep neural networks on the 69 real-world, tabular classification datasets from the OpenML-CC18 benchmark, SCARF not only improves classification accuracy in the fully-supervised setting but does so also in the presence of label noise and in the semi-supervised setting where only a fraction of the available training data is labeled. We show that SCARF complements existing strategies and outperforms alternatives like autoencoders. We conduct comprehensive ablations, detailing the importance of a range of factors.
Forward citations
Cited by 7 Pith papers
-
The Importance of Encoder Choice:A Tabular-Image Study
Tabular encoder choice reorders multimodal rankings, can erase apparent fusion gains, and requires non-vanilla extraction for in-context learning models to avoid train-test representation shift.
-
TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning
TIME fuses frozen TabPFN tabular embeddings with image features and beats MLP and NCART baselines on five tabular-image datasets, including incomplete medical data.
-
iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
iStructTab reports that ordering tabular features via a graph-based descriptor score before transformer fusion improves multimodal image-table classification on most of six benchmarks, with uneven gains.
-
Self-Supervised Representations for Binary Program Clustering: From Empirical Study to Retrieval-Augmented Learning
VIME-R, a nearest-neighbor variant of VIME, reports state-of-the-art clustering homogeneity on the Ember and Bodmas malware datasets.
-
MIRRAMS: Learning Robust Tabular Models under Unseen Missingness Shifts
A training objective built on mutual-information robustness conditions plus extra random masking improves tabular model accuracy under missingness shifts between train and test, with gains also in fully observed settings.
-
(GG) MoE vs. MLP on Tabular Data
A Gumbel-Softmax-gated mixture of experts with numerical embeddings matches MLP accuracy on 38 tabular datasets while using roughly 10x fewer parameters, but its edge over MLP is not statistically significant.
-
Random at First, Fast at Last: NTK-Guided Fourier Pre-Processing for Tabular DL
Fixed random Fourier projections on tabular inputs are claimed to bound the NTK, speed up gradient descent, and improve accuracy across four architectures and eight benchmarks.
Discussion (0). Continue with ORCID to comment.