Pith. sign in

REVIEW 4 major objections 37 references

A single frozen transformer pretrained on early-fused bank event sequences beats task-specific feature models and lifts production NPV by about 1 percent.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:19 UTC pith:P4CVKVTF

load-bearing objection Solid industrial recipe with real A/B NPV evidence; novelty is integration under bank constraints, not a new modeling principle. the 4 major comments →

arxiv 2607.09955 v1 pith:P4CVKVTF submitted 2026-07-10 cs.LG cs.AI

A Foundation Model for Multimodal Event Sequences in Financial Applications

classification cs.LG cs.AI
keywords foundation modelsmultimodal event sequencesself-supervised learningtransactional datafinancial applicationsnext-event predictionearly fusionuser embeddings
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Banks usually train a separate model for each prediction job on hand-built tabular features, which wastes effort and leaves most raw transaction, click, and communication history unused. This paper shows that unifying those heterogeneous events into one chronological sequence, pretraining a transformer with next-event prediction, and freezing the resulting user embeddings produces reusable representations. Lightweight neural heads that combine those embeddings with the existing engineered features then outperform the old task-specific baselines on multiple customer-response tasks. The same system, once deployed, improved net present value of communications by roughly one percent in live A/B tests while cutting model-development overhead.

Core claim

A foundation transformer pretrained with next-event prediction on early-fused multimodal sequences of user events (transactions, clickstream, communications) yields general-purpose user embeddings that, when frozen and concatenated with existing engineered features and passed through lightweight tabular heads, outperform traditional task-specific feature-based models offline and deliver measurable NPV gains in production at a major bank.

What carries the argument

Early-fusion next-event-prediction transformer: events from all modalities are merged into one chronological sequence, each heterogeneous event is turned into a fixed embedding by attribute-level self-attention and pooling, a causal GPT-style backbone models the sequence, and the mean-pooled final hidden states become the frozen user embedding z_seq that is later concatenated with a tabular embedding of engineered features.

Load-bearing premise

Mean-pooled hidden states from a frozen next-event-pretrained causal transformer already contain enough task-independent behavioral information that lightweight heads on top of the embedding plus engineered features can solve diverse downstream objectives without ever fine-tuning the sequence backbone.

What would settle it

Train the identical architecture end-to-end supervised on each downstream label (or fine-tune the backbone) and check whether the frozen-embedding pipeline still matches or beats it on held-out ROC AUC and on live NPV; if supervised or fine-tuned versions pull substantially ahead, the claim that frozen next-event embeddings are sufficient general-purpose representations fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • One pretrained backbone can be reused frozen across many banking prediction tasks, cutting the cost of building and maintaining separate sequence models.
  • Early fusion of transactions, clicks and communications yields higher average ROC AUC than modeling each modality separately and late-fusing.
  • Larger model capacity and longer context continue to improve offline metrics, so further scale is expected to help until inference cost becomes the binding constraint.
  • Production scoring pipelines can replace ensembles of hand-crafted GBDT models with the hybrid embedding-plus-features system and still satisfy business constraints while raising total NPV.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same early-fusion next-event objective could be applied to any multi-source event log (insurance claims, retail loyalty, telecom) where engineered tabular features already exist and must be preserved.
  • If the frozen embedding truly discards only redundant temporal detail, then periodic lightweight adapter updates without full backbone retraining may keep the system current as new products appear.
  • The modest absolute ROC-AUC lifts that still produce 1 percent NPV gains suggest that ranking quality at the very top of the score distribution, not average discrimination, is what the optimizer actually monetizes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 0 minor

Summary. The paper proposes a foundation-model pipeline for financial user modeling: heterogeneous events (transactions, clickstream, communications) are early-fused into one chronological sequence, encoded by a pad+intra-event-attention Event Encoder and a causal GPT-style transformer, and pretrained with next-event prediction (CE on present categoricals + MSE on present numericals). The sequence backbone is frozen; mean-pooled user embeddings are adapted by a small MLP and concatenated with a TabNN embedding of existing engineered features, then used by lightweight task heads. Offline, on four binary product-response tasks, pretrained+features (42M) reaches avg ROC AUC 0.801 vs 0.792 for engineered features alone, with ablations on scale, context length, pretraining vs supervised, modalities, and early vs late fusion. Online A/B at a large bank reports ~+1% total NPV uplift under a joint contact optimizer, after which the system was productionized.

Significance. If the multi-task reuse and online gains hold under tighter causal isolation, this is a practically important industrial result: a single frozen multimodal sequence backbone plus existing tabular features can replace or improve many task-specific models while cutting development cost. Strengths include real production deployment, a coherent early-fusion design for heterogeneous events, and a reasonably complete ablation suite (scale, context, pretraining, modalities, fusion). The contribution is primarily systems/empirical rather than theoretical; its value for the community is a documented large-scale recipe and evidence that NEP pretraining on bank multimodal logs transfers to engagement scoring when combined with engineered features.

major comments (4)
  1. §6 Online Results: The reported +1% total NPV is not cleanly attributable to the frozen multimodal foundation model. Only 6 of >40 products use the new scores; scores enter an existing joint optimizer as score×NPV under shared channel capacity, and the metric is total NPV across all products. Under joint allocation, score changes on a high-NPV subset can reallocate exposure and revenue among untreated products, so Δ need not equal the incremental value of z_seq. Please report product-level and channel-level NPV (treated vs untreated), a holdout where the optimizer is fixed or non-joint, and/or an ablation that swaps only the scoring model while freezing allocation rules, so the foundation-model claim is isolated from optimizer reallocation.
  2. §5.2 Table 2 and §5.1.2: The multi-task foundation claim rests on four anonymized binary “customer response to communications” tasks with a small average offline lift (0.801 vs 0.792) and no standard errors, bootstrap CIs, or significance tests. Without uncertainty, it is unclear whether the 0.009 avg gain is stable across seeds/splits. Moreover, all tasks share the same objective family, so they do not stress the §3 goal of representations reusable across diverse business tasks (risk, fraud, etc.). Add uncertainty estimates and at least one qualitatively different task family, or narrow the claim to engagement/response scoring.
  3. §3 and §4.3 (frozen backbone / mean pooling): The design goal is that mean-pooled causal NEP states z_seq are task-independent enough that freezing Φ_core and training only TabNN+adapter+head suffices. The paper compares frozen pretrained vs supervised end-to-end (§5.4 Table 5) but does not ablate (i) fine-tuning the sequence backbone on downstream labels, (ii) alternative aggregators (last token, attention pooling, CLS), or (iii) whether task-critical temporal structure is lost under freeze+mean-pool. If fine-tuning or better pooling closes most of the gap, the “frozen foundation backbone” operational claim is weaker. Please add these ablations or qualify the reuse claim accordingly.
  4. §2 Related Work / §5 experiments: Prior financial sequence models (CoLES, NPPR, nuFormer, TREASURE) and industrial behavior FMs are discussed, but offline comparisons are only vs engineered features and internal fusion/pretraining variants—not vs strong sequential baselines (e.g., Txn-only NEP transformer with the same TabNN, CoLES-style contrastive embeddings, or late-fusion of modality-specific models beyond Table 7) under matched data and compute. Table 6’s Txn-only is helpful but incomplete. Without external sequential baselines on the same bank data, it is hard to credit early multimodal fusion and the Event Encoder specifically versus “any large sequence model + features.”

Circularity Check

0 steps flagged

No circularity: empirical pretrain–then–downstream evaluation; NEP objective is not algebraically identical to response labels or NPV.

full rationale

This is a standard industrial ML paper whose central claims rest on held-out ROC AUC comparisons (Tables 2–7) and an online A/B NPV uplift (§6), not on a first-principles derivation. The pretraining loss (next-event CE for categorical attributes + MSE for numerical attributes, §4.1.4) is defined on event attributes of the chronological multimodal sequence; downstream tasks are binary customer-response labels and business NPV under an existing optimizer. Those quantities are not the same object, so the reported lifts cannot reduce to the pretraining fit by construction. Design choices (NEP over CoLES/MLM/DeTPP, TabNN over DCNv2/TabM, early vs late fusion) are justified by internal ablations on proprietary data, not by a uniqueness theorem or self-citation that forbids alternatives. Related-work self-citations (e.g., Klenitskiy et al. on sequence autoencoding) appear only as background and are not load-bearing for the empirical claim. No fitted scalar is renamed as a prediction; no equation equates the foundation embedding to the downstream target. Circularity score is therefore 0.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

This is an empirical systems paper. The central claim rests on modeling choices (early fusion, NEP objective, frozen backbone, TabNN fusion) and industrial data assumptions rather than mathematical axioms. Free parameters are the usual deep-learning hyperparameters and architecture sizes chosen on proprietary validation. No new physical entities are postulated; the ‘Event Encoder’ and TabNN are architectural modules.

free parameters (5)
  • Model scale configs (2M/8M/42M: layers, heads, hidden size, context)
    Chosen under production inference constraints; default 8M, production-facing 42M selected as performance/cost trade-off (§5.1.3, §5.3).
  • Pretraining hyperparameters (lr=1e-3, wd=1e-2, dropout=0.15, batch/accum, epochs 2–10)
    Hand-set training recipe; checkpoint selected by validation loss (§5.1.4).
  • Downstream TabNN hyperparameters (d_emb=64, 3 layers, 4 heads, lr=0.004, 40 epochs)
    Fixed across experiments after internal benchmarking against FFNN, DCNv2, TabM (§4.4.2, §5.1.5).
  • Attribute embedding dim d and N_attr padding size
    Architectural free choices for heterogeneous event encoding (§4.1).
  • Loss composition (unweighted sum of CE over present categoricals + MSE over present numericals)
    Objective design choice; no learned task weights reported (§4.1.4).
axioms (5)
  • domain assumption Next-event prediction on chronological multimodal sequences yields transferable general-purpose user representations.
    Motivated by prior work and internal objective comparison (CoLES, MLM, DeTPP); treated as design premise in §4.4.1 and §3.
  • domain assumption Existing engineered tabular features remain complementary and should be retained rather than replaced.
    Stated in problem formulation §3; drives the frozen-seq + TabNN pipeline.
  • domain assumption Early fusion of heterogeneous modalities into one sequence is preferable to late fusion for cross-modal temporal interactions.
    Adopted as main design (§4); later ablated in §5.6 with modest average edge.
  • ad hoc to paper Mean pooling of causal transformer outputs is an adequate user embedding aggregator for diverse binary response tasks.
    Specified in §4.1.3 without comparison to last-token, attention pooling, or other aggregators.
  • standard math Standard transformer/GPT-2 inductive biases and AdamW optimization apply to bank event sequences.
    Background ML practice used throughout §4–5.
invented entities (2)
  • Event Encoder (pad + intra-event attribute self-attention + masked mean pool) no independent evidence
    purpose: Map heterogeneous multimodal events with different attribute sets into fixed-size event embeddings for a shared transformer.
    Architectural module introduced in §4.1.2; related to TP-BERTa-style intra-feature attention but specialized here. No independent evidence outside this system’s ablations.
  • TabNN (log1p+zscore numericals, feature-token transformer, learnable weighted aggregation, noise injection) no independent evidence
    purpose: Encode engineered user features into z_tab for concatenation with sequential embeddings.
    Selected after internal benchmarks; close to FT-Transformer with local tweaks (§4.2, §4.4.2). Not independently validated outside proprietary tasks.

pith-pipeline@v1.1.0-grok45 · 18397 in / 3583 out tokens · 38273 ms · 2026-07-14T14:19:27.216827+00:00 · methodology

0 comments
read the original abstract

Predictive modeling is a core component of modern financial services, where a wide range of tasks are traditionally addressed using separate models trained on manually engineered tabular features. This task-specific approach limits reuse and makes it difficult to fully exploit heterogeneous data sources such as transaction histories and digital interaction signals. In this paper, we present an approach based on pretraining a foundation transformer model on multimodal sequences of user events. Events from multiple data sources are unified into a single chronological sequence, enabling early fusion of heterogeneous modalities and learning of general-purpose representations via a next-event prediction objective. These representations are combined with existing engineered user features, on top of which lightweight neural models are trained for multiple downstream tasks. The proposed system outperforms traditional task-specific models while reducing development overhead. The approach was deployed in production at one of the biggest banks in Eastern Europe, resulting in measurable improvements in business metrics.

Figures

Figures reproduced from arXiv: 2607.09955 by Alexander Uglov, Alexey Vasilev, Anton Klenitskiy, Gleb Zaripov, Konstantin Zorin, Nikita Rusakov, Vladislav Meshkov.

Figure 1
Figure 1. Figure 1: Pretrained sequence model architecture. The sequential processing of heterogeneous user events (transactions, clicks, [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Schematic overview of the proposed downstream [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 10 linked inside Pith

  1. [1]

    Dmitrii Babaev, Nikita Ovsov, Ivan Kireev, Maria Ivanova, Gleb Gusev, Ivan Nazarov, and Alexander Tuzhilin. 2022. CoLES: Contrastive Learning for Event Sequences with Self-Supervision. InProceedings of the 2022 International Confer- ence on Management of Data. 1190–1199

  2. [2]

    Alexandra Bazarova, Maria Kovaleva, Ilya Kuleshov, Evgenia Romanenkova, Alexander Stepikin, Aleksandr Yugay, Dzhambulat Mollaev, Ivan Kireev, Andrey Savchenko, and Alexey Zaytsev. 2025. Learning Transactions Representations for Information Management in Banks: Mastering Local, Global, and External Knowledge.International Journal of Information Management ...

  3. [3]

    DT Braithwaite, Misael Cavalcanti, R Austin McEver, Hiroto Udagawa, Daniel Silva, Rohan Ramanath, Felipe Meneses, Arissa Yoshida, Evan Wingert, Matheus Ramos, et al. 2025. Your Spending Needs Attention: Modeling Financial Habits with Transformers.arXiv preprint arXiv:2507.23267(2025). A Foundation Model for Multimodal Event Sequences in Financial Applicat...

  4. [4]

    Xiangyi Chen, Kousik Rajesh, Matthew Lawhon, Zelun Wang, Hanyu Li, Haomiao Li, Saurabh Vishwas Joshi, Pong Eksombatchai, Jaewon Yang, Yi-Ping Hsu, et al

  5. [5]

    InProceedings of the Nineteenth ACM Conference on Recommender Systems

    PinFM: Foundation Model for User Activity Sequences at a Billion-scale Visual Discovery Platform. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 381–390

  6. [6]

    Jacek Dabrowski, Maria Janicka, Lukasz Sienkiewicz, Gergely Stomfai, Diet- mar Jannach, Francesco Barile, Marco Polignano, Claudio Pomo, and Abhishek Srivastava. 2025. RecSys Challenge 2025: Universal Behavioral Profiles for Rec- ommender Systems. InProceedings of the Nineteenth ACM Conference on Recom- mender Systems. 1389–1393

  7. [7]

    Huangliang Dai, Shixun Wu, Hairui Zhao, Jiajun Huang, Zizhe Jian, Yue Zhu, and Haiyang Hu. 2025. FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention. doi:10.48550/arXiv.2504.02211

  8. [8]

    Yingtong Dou, Zhimeng Jiang, Tianyi Zhang, Mingzhi Hu, Zhichao Xu, Shub- ham Jain, Uday Singh Saini, Xiran Fan, Jiarui Sun, Menghai Pan, et al . 2025. TransactionGPT.arXiv preprint arXiv:2511.08939(2025)

  9. [9]

    Jiahui Gong, Jingtao Ding, Fanjin Meng, Chen Yang, Hong Chen, Zuojian Wang, Haisheng Lu, and Yong Li. 2025. BehaveGPT: A Foundation Model for Large-scale User Behavior Modeling.arXiv preprint arXiv:2505.17631(2025)

  10. [10]

    Yury Gorishniy, Akim Kotelnikov, and Artem Babenko. 2025. TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/ forum?id=Sd4wYYOhmY

  11. [11]

    Yue Guo, Wentao Zhang, Xiaojun Zhang, Vincent W Zheng, and Yi Yang. 2025. Efficient Multi-Expert Tabular Language Model for Banking. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 2271–2281

  12. [12]

    Sanjeev Jha, Montserrat Guillen, and J Christopher Westland. 2012. Employing transaction aggregation strategy to detect credit card fraud.Expert systems with applications39, 16 (2012), 12650–12657

  13. [13]

    Ivan Karpukhin and Andrey Savchenko. 2024. Detecting the Future: All-at- Once Event Sequence Forecasting with Horizon Matching.arXiv preprint arXiv:2408.13131(2024)

  14. [14]

    Ivan Karpukhin and Andrey Savchenko. 2025. HT-Transformer: Event Sequences Classification by Accumulating Prefix Information with History Tokens.arXiv preprint arXiv:2508.01474(2025)

  15. [15]

    Kirill Khrylchenko, Artem Matveev, Sergei Makeev, and Vladimir Baikalov. 2025. Scaling Recommender Transformers to One Billion Parameters.arXiv preprint arXiv:2507.15994(2025)

  16. [16]

    Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter

  17. [17]

    arXiv:1706.02515 [cs.LG] https://arxiv

    Self-Normalizing Neural Networks. arXiv:1706.02515 [cs.LG] https://arxiv. org/abs/1706.02515

  18. [18]

    Anton Klenitskiy, Artem Fatkulin, Daria Denisova, Anton Pembek, and Alexey Vasilev. 2025. Encode Me If You Can: Learning Universal User Representations via Event Sequence Autoencoding. InProceedings of the Recommender Systems Challenge 2025. 26–30

  19. [19]

    Can Liu, Yuncong Gao, Li Sun, Jinghua Feng, Hao Yang, and Xiang Ao. 2022. User Behavior Pre-training for Online Fraud Detection. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3357–3365

  20. [20]

    Wenhan Lyu, Devashish Tyagi, Yihang Yang, Ziwei Li, Ajay Somani, Karthikeyan Shanmugasundaram, Nikola Andrejevic, Ferdi Adeputra, Curtis Zeng, Arun K Singh, et al. 2025. DV365: Extremely Long User History Modeling at Instagram. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 4717–4727

  21. [21]

    Sergei Makeev, Alexandr Andreev, Vladimir Baikalov, Vladislav Tytskiy, Aleksei Krasilnikov, and Kirill Khrylchenko. 2025. Blending Sequential Embeddings, Graphs, and Engineered Features: 4th Place Solution in RecSys Challenge 2025. InProceedings of the Recommender Systems Challenge 2025. 21–25

  22. [22]

    Dzhambulat Mollaev, Ivan Kireev, Mikhail Orlov, Alexander Kostin, Ivan Karpukhin, Maria Postnova, Gleb Gusev, and Andrey Savchenko. 2025. Multi- modal Banking Dataset: Understanding Client Needs through Event Sequences. InProceedings of the 34th ACM International Conference on Information and Knowl- edge Management. 6476–6480

  23. [23]

    Viktor Moskvoretskii, Dmitry Osin, Egor Shvetsov, Igor Udovichenko, Maxim Zhelnin, Andrey Dukhovny, Anna Zhimerikina, and Evgeny Burnaev. 2024. MLEM: Generative and Contrastive Learning as Distinct Modalities for Event Sequences.arXiv preprint arXiv:2401.15935(2024)

  24. [24]

    Maxim Ostroukhov, Ruslan Mikhailov, Vladimir Iashin, Artem Sokolov, Andrei Akshonov, Vitaly Protasov, Dmitrii Beloborodov, Vince Mullin, Roman Yokunda Enzmann, Georgios Kolovos, et al. 2026. PRAGMA: Revolut Foundation Model. arXiv preprint arXiv:2604.08649(2026)

  25. [25]

    Nikil Pancha, Andrew Zhai, Jure Leskovec, and Charles Rosenberg. 2022. Pinner- Former: Sequence Modeling for User Representation at Pinterest. InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. 3702–3712

  26. [26]

    Gustavo Polleti, Marlesson Santana, and Eduardo Fontes. 2025. Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions.arXiv preprint arXiv:2511.12154(2025)

  27. [27]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners.OpenAI (2019). https://cdn.openai.com/better-language-models/language_models_are_ unsupervised_multitask_learners.pdf Accessed: 2024-11-15

  28. [28]

    Natraj Raman, Sumitra Ganesh, and Manuela Veloso. 2024. Scalable Rep- resentation Learning for Multimodal Tabular Transactions.arXiv preprint arXiv:2410.07851(2024)

  29. [29]

    Artem Sakhno, Ivan Kireev, Dmitrii Babaev, Maxim Savchenko, Gleb Gusev, and Andrey Savchenko. 2025. PyTorch-Lifestream: Learning Embeddings on Discrete Event Sequences. InProceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. 11104–11108

  30. [30]

    Yuki Sawada, Rintaro Hasegawa, Yuhi Nagatsuma, Shugo Takei, Kazuhito Yonekawa, and Hiromu Auchi. 2025. Toward Universal User Representations: Contrastive Learning with Transformers and Embedding Ensembles. InProceed- ings of the Recommender Systems Challenge 2025. 51–55

  31. [31]

    Aleksei Shestov, Omar Zoloev, Maksim Makarenko, Mikhail Orlov, Egor Fadeev, Ivan Kireev, and Andrey Savchenko. 2025. LLM4ES: Learning user embeddings from event sequences via large language models. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 5238–5242

  32. [32]

    Piotr Skalski, David Sutton, Stuart Burrell, Iker Perez, and Jason Wong. 2023. To- wards a Foundation Purchasing Model: Pretrained Generative Autoregression on Transaction Sequences. InProceedings of the Fourth ACM International Conference on AI in Finance. 141–149

  33. [33]

    Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. 1785–1797. doi:10.1145/3442381.3450078

  34. [34]

    Ziming Wang, Qianru Wu, Baolin Zheng, Junjie Wang, Kaiyu Huang, and Yanjie Shi. 2023. Sequence As Genes: An User Behavior Modeling Framework for Fraud Transaction Detection in E-commerce. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 5194–5203

  35. [35]

    Xue Xia, Pong Eksombatchai, Nikil Pancha, Dhruvil Deven Badani, Po-Wei Wang, Neng Gu, Saurabh Vishwas Joshi, Nazanin Farahpour, Zhiyuan Zhang, and An- drew Zhai. 2023. TransAct: Transformer-based Realtime User Action Model for Recommendation at Pinterest. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 5249–5259

  36. [36]

    Chen, Jimeng Sun, Jian Wu, and Jintai Chen

    Jiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu, Danny Z. Chen, Jimeng Sun, Jian Wu, and Jintai Chen. 2024. Making Pre-trained Language Models Great on Tabular Prediction. arXiv:2403.01841 [cs.CL] https://arxiv.org/abs/2403.01841

  37. [37]

    Chin-Chia Michael Yeh, Uday Singh Saini, Xin Dai, Xiran Fan, Shubham Jain, Yujie Fan, Jiarui Sun, Junpeng Wang, Menghai Pan, Yingtong Dou, et al. 2025. TREASURE: A Transformer-Based Foundation Model for High-Volume Transac- tion Understanding.arXiv preprint arXiv:2511.19693(2025)