REVIEW 21 cited by
SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Tabular data underpins numerous high-impact applications of machine learning from fraud detection to genomics and healthcare. Classical approaches to solving tabular problems, such as gradient boosting and random forests, are widely used by practitioners. However, recent deep learning methods have achieved a degree of performance competitive with popular techniques. We devise a hybrid deep learning approach to solving tabular data problems. Our method, SAINT, performs attention over both rows and columns, and it includes an enhanced embedding method. We also study a new contrastive self-supervised pre-training method for use when labels are scarce. SAINT consistently improves performance over previous deep learning methods, and it even outperforms gradient boosting methods, including XGBoost, CatBoost, and LightGBM, on average over a variety of benchmark tasks.
Forward citations
Cited by 21 Pith papers
-
Adaptive Intrusion Detection System using Transformer-Based Neural Networks and Continual Learning Approach with Adversarial Investigation
Benign-anchored class-balanced replay with a tabular transformer nearly matches joint training on CICIDS2017, and buffer poisoning attacks show the replay store must be treated as security-critical state.
-
Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions
Sparse Bayesian-network factorization plus sparsity-aware regression yields polynomial TV rates for high-dimensional mixed-type distribution estimation, beating classical histogram rates under sparsity.
-
Foundation Models for Credit Risk Prediction: A Game Changer?
Tabular foundation models, used zero-shot, match or beat tuned gradient boosting on average in credit PD and LGD benchmarks, with a larger edge on small datasets.
-
No Need to Train Your RDB Foundation Model
Column-wise, parameter-free JUICE encodings let single-table ICL models solve multi-table RDB prediction tasks with no training or fine-tuning.
-
LakeMLB: Data Lake Machine Learning Benchmark
LakeMLB is a new six-dataset benchmark for multi-table machine learning in data lakes; experiments find pretraining helps in Union scenarios and feature augmentation helps in Join scenarios.
-
City-Level Foreign Direct Investment Prediction with Tabular Learning on Judicial Data
A tabular learning model trained on 380 judicial-performance indicators from 12 million court documents predicts Chinese city-level FDI with R2 up to 0.92 on mixed-year and cross-time benchmarks.
-
Constructive Universal Approximation and Sure Convergence for Multi-Layer Neural Networks
A multi-layer network of sparse indicator neurons, trained by greedy boosting with random candidate search, is shown to be a universal approximator and to converge to a sample-optimal model eventually.
-
LAKEGEN: A LLM-based Tabular Corpus Generator for Evaluating Dataset Discovery in Data Lakes
LAKEGEN builds synthetic, domain-specific tabular benchmarks using ontologies and an LLM, and shows current dataset discovery methods struggle on the resulting semantic joinability tasks.
-
On Finetuning Tabular Foundation Models
Full finetuning of TabPFNv2 outperforms in-context learning and partial finetuning on medium tabular datasets, and its gains come from sharper query-key attention that better reflects target similarity.
-
TabFlex: Scaling Tabular Learning to Millions with Linear Attention
Linear attention lets a TabPFN-style model process millions of tabular samples in seconds with near-identical accuracy on small datasets.
-
TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning
TIME fuses frozen TabPFN tabular embeddings with image features and beats MLP and NCART baselines on five tabular-image datasets, including incomplete medical data.
-
When Shift Happens - Confounding Is to Blame
Under hidden confounding shifts, predictive information reduces to conditional informativeness minus a residual, a result the authors use to explain ERM's surprising OOD competitiveness and the value of all-covariate models.
-
MultiTab: A Comprehensive Benchmark Suite for Multi-Dimensional Evaluation in Tabular Domains
A regime-stratified benchmark of 196 tabular datasets shows that model rankings depend strongly on dataset characteristics such as sample size, feature correlation, and label imbalance.
-
Pattern-Aware Graph Neural Networks for Handling Missing Data
Explicitly encoding missingness patterns in bipartite GNNs yields ~17% average balanced-accuracy gains over GRAPE on seven UCI datasets with natural missingness, with random embeddings nearly matching learned ones.
-
Modeling Decisions in Blockchain Analytics: A Leakage-Aware Evaluation of Tree-Based vs. Sequential Models
The paper reports that XGBoost beats Transformer and BiLSTM models for Ethereum actor classification after masking certain high-signal contracts, and that sequence order adds little signal.
-
Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models
All nine tested tabular foundation models degrade under distribution shift, and real-world pretraining provides no robustness advantage over synthetic pretraining.
-
TabLoRA: Parameter-Efficient Low-Rank Ensemble Learning for Large-Scale Tabular Data
Shared-backbone low-rank ensemble adapters let neural tabular models match much of full-ensemble accuracy on large data without linear parameter growth or frequent OOMs.
-
PECKER: A Precisely Efficient Critical Knowledge Erasure Recipe For Machine Unlearning in Diffusion Models
Spline numerical encodings can match or beat standard scaling on tabular nets, but PLE is most robust for classification and learnable knots add substantial training cost.
-
MIRRAMS: Learning Robust Tabular Models under Unseen Missingness Shifts
A training objective built on mutual-information robustness conditions plus extra random masking improves tabular model accuracy under missingness shifts between train and test, with gains also in fully observed settings.
-
Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification
Across 12 fixed-hyperparameter tabular benchmarks, KANs beat MLPs on 9/12 datasets in accuracy and 10/12 in F1, yet cost ~16x parameters; the paper's significance tests are internally inconsistent.
-
Random at First, Fast at Last: NTK-Guided Fourier Pre-Processing for Tabular DL
Fixed random Fourier projections on tabular inputs are claimed to bound the NTK, speed up gradient descent, and improve accuracy across four architectures and eight benchmarks.
Discussion (0). Continue with ORCID to comment.