Pith. sign in

REVIEW 21 cited by

SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.01342 v1 pith:ZWYL2XTT submitted 2021-06-02 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningtabulardatadeepmethodmethodssaintattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tabular data underpins numerous high-impact applications of machine learning from fraud detection to genomics and healthcare. Classical approaches to solving tabular problems, such as gradient boosting and random forests, are widely used by practitioners. However, recent deep learning methods have achieved a degree of performance competitive with popular techniques. We devise a hybrid deep learning approach to solving tabular data problems. Our method, SAINT, performs attention over both rows and columns, and it includes an enhanced embedding method. We also study a new contrastive self-supervised pre-training method for use when labels are scarce. SAINT consistently improves performance over previous deep learning methods, and it even outperforms gradient boosting methods, including XGBoost, CatBoost, and LightGBM, on average over a variety of benchmark tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 120 citations worldwide. Full citation record

  1. Adaptive Intrusion Detection System using Transformer-Based Neural Networks and Continual Learning Approach with Adversarial Investigation

    cs.CR 2026-08 conditional novelty 6.0 of 10

    Benign-anchored class-balanced replay with a tabular transformer nearly matches joint training on CICIDS2017, and buffer poisoning attacks show the replay store must be treated as security-critical state.

  2. Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions

    stat.ML 2026-07 conditional novelty 6.0 of 10

    Sparse Bayesian-network factorization plus sparsity-aware regression yields polynomial TV rates for high-dimensional mixed-type distribution estimation, beating classical histogram rates under sparsity.

  3. Foundation Models for Credit Risk Prediction: A Game Changer?

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Tabular foundation models, used zero-shot, match or beat tuned gradient boosting on average in credit PD and LGD benchmarks, with a larger edge on small datasets.

  4. No Need to Train Your RDB Foundation Model

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Column-wise, parameter-free JUICE encodings let single-table ICL models solve multi-table RDB prediction tasks with no training or fine-tuning.

  5. LakeMLB: Data Lake Machine Learning Benchmark

    cs.LG 2026-02 conditional novelty 6.0 of 10

    LakeMLB is a new six-dataset benchmark for multi-table machine learning in data lakes; experiments find pretraining helps in Union scenarios and feature augmentation helps in Join scenarios.

  6. City-Level Foreign Direct Investment Prediction with Tabular Learning on Judicial Data

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A tabular learning model trained on 380 judicial-performance indicators from 12 million court documents predicts Chinese city-level FDI with R2 up to 0.92 on mixed-year and cross-time benchmarks.

  7. Constructive Universal Approximation and Sure Convergence for Multi-Layer Neural Networks

    stat.ML 2025-07 conditional novelty 6.0 of 10

    A multi-layer network of sparse indicator neurons, trained by greedy boosting with random candidate search, is shown to be a universal approximator and to converge to a sample-optimal model eventually.

  8. LAKEGEN: A LLM-based Tabular Corpus Generator for Evaluating Dataset Discovery in Data Lakes

    cs.DB 2025-07 conditional novelty 6.0 of 10

    LAKEGEN builds synthetic, domain-specific tabular benchmarks using ontologies and an LLM, and shows current dataset discovery methods struggle on the resulting semantic joinability tasks.

  9. On Finetuning Tabular Foundation Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Full finetuning of TabPFNv2 outperforms in-context learning and partial finetuning on medium tabular datasets, and its gains come from sharper query-key attention that better reflects target similarity.

  10. TabFlex: Scaling Tabular Learning to Millions with Linear Attention

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Linear attention lets a TabPFN-style model process millions of tabular samples in seconds with near-identical accuracy on small datasets.

  11. TIME: TabPFN-Integrated Multimodal Engine for Robust Tabular-Image Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TIME fuses frozen TabPFN tabular embeddings with image features and beats MLP and NCART baselines on five tabular-image datasets, including incomplete medical data.

  12. When Shift Happens - Confounding Is to Blame

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Under hidden confounding shifts, predictive information reduces to conditional informativeness minus a residual, a result the authors use to explain ERM's surprising OOD competitiveness and the value of all-covariate models.

  13. MultiTab: A Comprehensive Benchmark Suite for Multi-Dimensional Evaluation in Tabular Domains

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A regime-stratified benchmark of 196 tabular datasets shows that model rankings depend strongly on dataset characteristics such as sample size, feature correlation, and label imbalance.

  14. Pattern-Aware Graph Neural Networks for Handling Missing Data

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Explicitly encoding missingness patterns in bipartite GNNs yields ~17% average balanced-accuracy gains over GRAPE on seven UCI datasets with natural missingness, with random embeddings nearly matching learned ones.

  15. Modeling Decisions in Blockchain Analytics: A Leakage-Aware Evaluation of Tree-Based vs. Sequential Models

    cs.LG 2026-07 reject novelty 5.0 of 10

    The paper reports that XGBoost beats Transformer and BiLSTM models for Ethereum actor classification after masking certain high-signal contracts, and that sequence order adds little signal.

  16. Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

    cs.LG 2026-07 conditional novelty 5.0 of 10

    All nine tested tabular foundation models degrade under distribution shift, and real-world pretraining provides no robustness advantage over synthetic pretraining.

  17. TabLoRA: Parameter-Efficient Low-Rank Ensemble Learning for Large-Scale Tabular Data

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Shared-backbone low-rank ensemble adapters let neural tabular models match much of full-ensemble accuracy on large data without linear parameter growth or frequent OOMs.

  18. PECKER: A Precisely Efficient Critical Knowledge Erasure Recipe For Machine Unlearning in Diffusion Models

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Spline numerical encodings can match or beat standard scaling on tabular nets, but PLE is most robust for classification and learnable knots add substantial training cost.

  19. MIRRAMS: Learning Robust Tabular Models under Unseen Missingness Shifts

    stat.ML 2025-07 conditional novelty 5.0 of 10

    A training objective built on mutual-information robustness conditions plus extra random masking improves tabular model accuracy under missingness shifts between train and test, with gains also in fully observed settings.

  20. Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification

    cs.LG 2026-07 reject novelty 4.0 of 10

    Across 12 fixed-hyperparameter tabular benchmarks, KANs beat MLPs on 9/12 datasets in accuracy and 10/12 in F1, yet cost ~16x parameters; the paper's significance tests are internally inconsistent.

  21. Random at First, Fast at Last: NTK-Guided Fourier Pre-Processing for Tabular DL

    cs.LG 2025-06 conditional novelty 3.0 of 10

    Fixed random Fourier projections on tabular inputs are claimed to bound the NTK, speed up gradient descent, and improve accuracy across four architectures and eight benchmarks.

Pith tools