REVIEW 4 major objections 37 references
A single frozen transformer pretrained on early-fused bank event sequences beats task-specific feature models and lifts production NPV by about 1 percent.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 14:19 UTC pith:P4CVKVTF
load-bearing objection Solid industrial recipe with real A/B NPV evidence; novelty is integration under bank constraints, not a new modeling principle. the 4 major comments →
A Foundation Model for Multimodal Event Sequences in Financial Applications
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A foundation transformer pretrained with next-event prediction on early-fused multimodal sequences of user events (transactions, clickstream, communications) yields general-purpose user embeddings that, when frozen and concatenated with existing engineered features and passed through lightweight tabular heads, outperform traditional task-specific feature-based models offline and deliver measurable NPV gains in production at a major bank.
What carries the argument
Early-fusion next-event-prediction transformer: events from all modalities are merged into one chronological sequence, each heterogeneous event is turned into a fixed embedding by attribute-level self-attention and pooling, a causal GPT-style backbone models the sequence, and the mean-pooled final hidden states become the frozen user embedding z_seq that is later concatenated with a tabular embedding of engineered features.
Load-bearing premise
Mean-pooled hidden states from a frozen next-event-pretrained causal transformer already contain enough task-independent behavioral information that lightweight heads on top of the embedding plus engineered features can solve diverse downstream objectives without ever fine-tuning the sequence backbone.
What would settle it
Train the identical architecture end-to-end supervised on each downstream label (or fine-tune the backbone) and check whether the frozen-embedding pipeline still matches or beats it on held-out ROC AUC and on live NPV; if supervised or fine-tuned versions pull substantially ahead, the claim that frozen next-event embeddings are sufficient general-purpose representations fails.
If this is right
- One pretrained backbone can be reused frozen across many banking prediction tasks, cutting the cost of building and maintaining separate sequence models.
- Early fusion of transactions, clicks and communications yields higher average ROC AUC than modeling each modality separately and late-fusing.
- Larger model capacity and longer context continue to improve offline metrics, so further scale is expected to help until inference cost becomes the binding constraint.
- Production scoring pipelines can replace ensembles of hand-crafted GBDT models with the hybrid embedding-plus-features system and still satisfy business constraints while raising total NPV.
Where Pith is reading between the lines
- The same early-fusion next-event objective could be applied to any multi-source event log (insurance claims, retail loyalty, telecom) where engineered tabular features already exist and must be preserved.
- If the frozen embedding truly discards only redundant temporal detail, then periodic lightweight adapter updates without full backbone retraining may keep the system current as new products appear.
- The modest absolute ROC-AUC lifts that still produce 1 percent NPV gains suggest that ranking quality at the very top of the score distribution, not average discrimination, is what the optimizer actually monetizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a foundation-model pipeline for financial user modeling: heterogeneous events (transactions, clickstream, communications) are early-fused into one chronological sequence, encoded by a pad+intra-event-attention Event Encoder and a causal GPT-style transformer, and pretrained with next-event prediction (CE on present categoricals + MSE on present numericals). The sequence backbone is frozen; mean-pooled user embeddings are adapted by a small MLP and concatenated with a TabNN embedding of existing engineered features, then used by lightweight task heads. Offline, on four binary product-response tasks, pretrained+features (42M) reaches avg ROC AUC 0.801 vs 0.792 for engineered features alone, with ablations on scale, context length, pretraining vs supervised, modalities, and early vs late fusion. Online A/B at a large bank reports ~+1% total NPV uplift under a joint contact optimizer, after which the system was productionized.
Significance. If the multi-task reuse and online gains hold under tighter causal isolation, this is a practically important industrial result: a single frozen multimodal sequence backbone plus existing tabular features can replace or improve many task-specific models while cutting development cost. Strengths include real production deployment, a coherent early-fusion design for heterogeneous events, and a reasonably complete ablation suite (scale, context, pretraining, modalities, fusion). The contribution is primarily systems/empirical rather than theoretical; its value for the community is a documented large-scale recipe and evidence that NEP pretraining on bank multimodal logs transfers to engagement scoring when combined with engineered features.
major comments (4)
- §6 Online Results: The reported +1% total NPV is not cleanly attributable to the frozen multimodal foundation model. Only 6 of >40 products use the new scores; scores enter an existing joint optimizer as score×NPV under shared channel capacity, and the metric is total NPV across all products. Under joint allocation, score changes on a high-NPV subset can reallocate exposure and revenue among untreated products, so Δ need not equal the incremental value of z_seq. Please report product-level and channel-level NPV (treated vs untreated), a holdout where the optimizer is fixed or non-joint, and/or an ablation that swaps only the scoring model while freezing allocation rules, so the foundation-model claim is isolated from optimizer reallocation.
- §5.2 Table 2 and §5.1.2: The multi-task foundation claim rests on four anonymized binary “customer response to communications” tasks with a small average offline lift (0.801 vs 0.792) and no standard errors, bootstrap CIs, or significance tests. Without uncertainty, it is unclear whether the 0.009 avg gain is stable across seeds/splits. Moreover, all tasks share the same objective family, so they do not stress the §3 goal of representations reusable across diverse business tasks (risk, fraud, etc.). Add uncertainty estimates and at least one qualitatively different task family, or narrow the claim to engagement/response scoring.
- §3 and §4.3 (frozen backbone / mean pooling): The design goal is that mean-pooled causal NEP states z_seq are task-independent enough that freezing Φ_core and training only TabNN+adapter+head suffices. The paper compares frozen pretrained vs supervised end-to-end (§5.4 Table 5) but does not ablate (i) fine-tuning the sequence backbone on downstream labels, (ii) alternative aggregators (last token, attention pooling, CLS), or (iii) whether task-critical temporal structure is lost under freeze+mean-pool. If fine-tuning or better pooling closes most of the gap, the “frozen foundation backbone” operational claim is weaker. Please add these ablations or qualify the reuse claim accordingly.
- §2 Related Work / §5 experiments: Prior financial sequence models (CoLES, NPPR, nuFormer, TREASURE) and industrial behavior FMs are discussed, but offline comparisons are only vs engineered features and internal fusion/pretraining variants—not vs strong sequential baselines (e.g., Txn-only NEP transformer with the same TabNN, CoLES-style contrastive embeddings, or late-fusion of modality-specific models beyond Table 7) under matched data and compute. Table 6’s Txn-only is helpful but incomplete. Without external sequential baselines on the same bank data, it is hard to credit early multimodal fusion and the Event Encoder specifically versus “any large sequence model + features.”
Circularity Check
No circularity: empirical pretrain–then–downstream evaluation; NEP objective is not algebraically identical to response labels or NPV.
full rationale
This is a standard industrial ML paper whose central claims rest on held-out ROC AUC comparisons (Tables 2–7) and an online A/B NPV uplift (§6), not on a first-principles derivation. The pretraining loss (next-event CE for categorical attributes + MSE for numerical attributes, §4.1.4) is defined on event attributes of the chronological multimodal sequence; downstream tasks are binary customer-response labels and business NPV under an existing optimizer. Those quantities are not the same object, so the reported lifts cannot reduce to the pretraining fit by construction. Design choices (NEP over CoLES/MLM/DeTPP, TabNN over DCNv2/TabM, early vs late fusion) are justified by internal ablations on proprietary data, not by a uniqueness theorem or self-citation that forbids alternatives. Related-work self-citations (e.g., Klenitskiy et al. on sequence autoencoding) appear only as background and are not load-bearing for the empirical claim. No fitted scalar is renamed as a prediction; no equation equates the foundation embedding to the downstream target. Circularity score is therefore 0.
Axiom & Free-Parameter Ledger
free parameters (5)
- Model scale configs (2M/8M/42M: layers, heads, hidden size, context)
- Pretraining hyperparameters (lr=1e-3, wd=1e-2, dropout=0.15, batch/accum, epochs 2–10)
- Downstream TabNN hyperparameters (d_emb=64, 3 layers, 4 heads, lr=0.004, 40 epochs)
- Attribute embedding dim d and N_attr padding size
- Loss composition (unweighted sum of CE over present categoricals + MSE over present numericals)
axioms (5)
- domain assumption Next-event prediction on chronological multimodal sequences yields transferable general-purpose user representations.
- domain assumption Existing engineered tabular features remain complementary and should be retained rather than replaced.
- domain assumption Early fusion of heterogeneous modalities into one sequence is preferable to late fusion for cross-modal temporal interactions.
- ad hoc to paper Mean pooling of causal transformer outputs is an adequate user embedding aggregator for diverse binary response tasks.
- standard math Standard transformer/GPT-2 inductive biases and AdamW optimization apply to bank event sequences.
invented entities (2)
-
Event Encoder (pad + intra-event attribute self-attention + masked mean pool)
no independent evidence
-
TabNN (log1p+zscore numericals, feature-token transformer, learnable weighted aggregation, noise injection)
no independent evidence
read the original abstract
Predictive modeling is a core component of modern financial services, where a wide range of tasks are traditionally addressed using separate models trained on manually engineered tabular features. This task-specific approach limits reuse and makes it difficult to fully exploit heterogeneous data sources such as transaction histories and digital interaction signals. In this paper, we present an approach based on pretraining a foundation transformer model on multimodal sequences of user events. Events from multiple data sources are unified into a single chronological sequence, enabling early fusion of heterogeneous modalities and learning of general-purpose representations via a next-event prediction objective. These representations are combined with existing engineered user features, on top of which lightweight neural models are trained for multiple downstream tasks. The proposed system outperforms traditional task-specific models while reducing development overhead. The approach was deployed in production at one of the biggest banks in Eastern Europe, resulting in measurable improvements in business metrics.
Figures
Reference graph
Works this paper leans on
-
[1]
Dmitrii Babaev, Nikita Ovsov, Ivan Kireev, Maria Ivanova, Gleb Gusev, Ivan Nazarov, and Alexander Tuzhilin. 2022. CoLES: Contrastive Learning for Event Sequences with Self-Supervision. InProceedings of the 2022 International Confer- ence on Management of Data. 1190–1199
2022
-
[2]
Alexandra Bazarova, Maria Kovaleva, Ilya Kuleshov, Evgenia Romanenkova, Alexander Stepikin, Aleksandr Yugay, Dzhambulat Mollaev, Ivan Kireev, Andrey Savchenko, and Alexey Zaytsev. 2025. Learning Transactions Representations for Information Management in Banks: Mastering Local, Global, and External Knowledge.International Journal of Information Management ...
2025
-
[3]
DT Braithwaite, Misael Cavalcanti, R Austin McEver, Hiroto Udagawa, Daniel Silva, Rohan Ramanath, Felipe Meneses, Arissa Yoshida, Evan Wingert, Matheus Ramos, et al. 2025. Your Spending Needs Attention: Modeling Financial Habits with Transformers.arXiv preprint arXiv:2507.23267(2025). A Foundation Model for Multimodal Event Sequences in Financial Applicat...
Pith/arXiv arXiv 2025
-
[4]
Xiangyi Chen, Kousik Rajesh, Matthew Lawhon, Zelun Wang, Hanyu Li, Haomiao Li, Saurabh Vishwas Joshi, Pong Eksombatchai, Jaewon Yang, Yi-Ping Hsu, et al
-
[5]
InProceedings of the Nineteenth ACM Conference on Recommender Systems
PinFM: Foundation Model for User Activity Sequences at a Billion-scale Visual Discovery Platform. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 381–390
-
[6]
Jacek Dabrowski, Maria Janicka, Lukasz Sienkiewicz, Gergely Stomfai, Diet- mar Jannach, Francesco Barile, Marco Polignano, Claudio Pomo, and Abhishek Srivastava. 2025. RecSys Challenge 2025: Universal Behavioral Profiles for Rec- ommender Systems. InProceedings of the Nineteenth ACM Conference on Recom- mender Systems. 1389–1393
2025
-
[7]
Huangliang Dai, Shixun Wu, Hairui Zhao, Jiajun Huang, Zizhe Jian, Yue Zhu, and Haiyang Hu. 2025. FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention. doi:10.48550/arXiv.2504.02211
-
[8]
Yingtong Dou, Zhimeng Jiang, Tianyi Zhang, Mingzhi Hu, Zhichao Xu, Shub- ham Jain, Uday Singh Saini, Xiran Fan, Jiarui Sun, Menghai Pan, et al . 2025. TransactionGPT.arXiv preprint arXiv:2511.08939(2025)
arXiv 2025
-
[9]
Jiahui Gong, Jingtao Ding, Fanjin Meng, Chen Yang, Hong Chen, Zuojian Wang, Haisheng Lu, and Yong Li. 2025. BehaveGPT: A Foundation Model for Large-scale User Behavior Modeling.arXiv preprint arXiv:2505.17631(2025)
Pith/arXiv arXiv 2025
-
[10]
Yury Gorishniy, Akim Kotelnikov, and Artem Babenko. 2025. TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/ forum?id=Sd4wYYOhmY
2025
-
[11]
Yue Guo, Wentao Zhang, Xiaojun Zhang, Vincent W Zheng, and Yi Yang. 2025. Efficient Multi-Expert Tabular Language Model for Banking. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 2271–2281
2025
-
[12]
Sanjeev Jha, Montserrat Guillen, and J Christopher Westland. 2012. Employing transaction aggregation strategy to detect credit card fraud.Expert systems with applications39, 16 (2012), 12650–12657
2012
-
[13]
Ivan Karpukhin and Andrey Savchenko. 2024. Detecting the Future: All-at- Once Event Sequence Forecasting with Horizon Matching.arXiv preprint arXiv:2408.13131(2024)
arXiv 2024
-
[14]
Ivan Karpukhin and Andrey Savchenko. 2025. HT-Transformer: Event Sequences Classification by Accumulating Prefix Information with History Tokens.arXiv preprint arXiv:2508.01474(2025)
Pith/arXiv arXiv 2025
-
[15]
Kirill Khrylchenko, Artem Matveev, Sergei Makeev, and Vladimir Baikalov. 2025. Scaling Recommender Transformers to One Billion Parameters.arXiv preprint arXiv:2507.15994(2025)
arXiv 2025
-
[16]
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter
-
[17]
arXiv:1706.02515 [cs.LG] https://arxiv
Self-Normalizing Neural Networks. arXiv:1706.02515 [cs.LG] https://arxiv. org/abs/1706.02515
-
[18]
Anton Klenitskiy, Artem Fatkulin, Daria Denisova, Anton Pembek, and Alexey Vasilev. 2025. Encode Me If You Can: Learning Universal User Representations via Event Sequence Autoencoding. InProceedings of the Recommender Systems Challenge 2025. 26–30
2025
-
[19]
Can Liu, Yuncong Gao, Li Sun, Jinghua Feng, Hao Yang, and Xiang Ao. 2022. User Behavior Pre-training for Online Fraud Detection. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3357–3365
2022
-
[20]
Wenhan Lyu, Devashish Tyagi, Yihang Yang, Ziwei Li, Ajay Somani, Karthikeyan Shanmugasundaram, Nikola Andrejevic, Ferdi Adeputra, Curtis Zeng, Arun K Singh, et al. 2025. DV365: Extremely Long User History Modeling at Instagram. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 4717–4727
2025
-
[21]
Sergei Makeev, Alexandr Andreev, Vladimir Baikalov, Vladislav Tytskiy, Aleksei Krasilnikov, and Kirill Khrylchenko. 2025. Blending Sequential Embeddings, Graphs, and Engineered Features: 4th Place Solution in RecSys Challenge 2025. InProceedings of the Recommender Systems Challenge 2025. 21–25
2025
-
[22]
Dzhambulat Mollaev, Ivan Kireev, Mikhail Orlov, Alexander Kostin, Ivan Karpukhin, Maria Postnova, Gleb Gusev, and Andrey Savchenko. 2025. Multi- modal Banking Dataset: Understanding Client Needs through Event Sequences. InProceedings of the 34th ACM International Conference on Information and Knowl- edge Management. 6476–6480
2025
-
[23]
Viktor Moskvoretskii, Dmitry Osin, Egor Shvetsov, Igor Udovichenko, Maxim Zhelnin, Andrey Dukhovny, Anna Zhimerikina, and Evgeny Burnaev. 2024. MLEM: Generative and Contrastive Learning as Distinct Modalities for Event Sequences.arXiv preprint arXiv:2401.15935(2024)
Pith/arXiv arXiv 2024
-
[24]
Maxim Ostroukhov, Ruslan Mikhailov, Vladimir Iashin, Artem Sokolov, Andrei Akshonov, Vitaly Protasov, Dmitrii Beloborodov, Vince Mullin, Roman Yokunda Enzmann, Georgios Kolovos, et al. 2026. PRAGMA: Revolut Foundation Model. arXiv preprint arXiv:2604.08649(2026)
Pith/arXiv arXiv 2026
-
[25]
Nikil Pancha, Andrew Zhai, Jure Leskovec, and Charles Rosenberg. 2022. Pinner- Former: Sequence Modeling for User Representation at Pinterest. InProceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. 3702–3712
2022
-
[26]
Gustavo Polleti, Marlesson Santana, and Eduardo Fontes. 2025. Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions.arXiv preprint arXiv:2511.12154(2025)
arXiv 2025
-
[27]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners.OpenAI (2019). https://cdn.openai.com/better-language-models/language_models_are_ unsupervised_multitask_learners.pdf Accessed: 2024-11-15
2019
-
[28]
Natraj Raman, Sumitra Ganesh, and Manuela Veloso. 2024. Scalable Rep- resentation Learning for Multimodal Tabular Transactions.arXiv preprint arXiv:2410.07851(2024)
Pith/arXiv arXiv 2024
-
[29]
Artem Sakhno, Ivan Kireev, Dmitrii Babaev, Maxim Savchenko, Gleb Gusev, and Andrey Savchenko. 2025. PyTorch-Lifestream: Learning Embeddings on Discrete Event Sequences. InProceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. 11104–11108
2025
-
[30]
Yuki Sawada, Rintaro Hasegawa, Yuhi Nagatsuma, Shugo Takei, Kazuhito Yonekawa, and Hiromu Auchi. 2025. Toward Universal User Representations: Contrastive Learning with Transformers and Embedding Ensembles. InProceed- ings of the Recommender Systems Challenge 2025. 51–55
2025
-
[31]
Aleksei Shestov, Omar Zoloev, Maksim Makarenko, Mikhail Orlov, Egor Fadeev, Ivan Kireev, and Andrey Savchenko. 2025. LLM4ES: Learning user embeddings from event sequences via large language models. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 5238–5242
2025
-
[32]
Piotr Skalski, David Sutton, Stuart Burrell, Iker Perez, and Jason Wong. 2023. To- wards a Foundation Purchasing Model: Pretrained Generative Autoregression on Transaction Sequences. InProceedings of the Fourth ACM International Conference on AI in Finance. 141–149
2023
-
[33]
Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. 1785–1797. doi:10.1145/3442381.3450078
-
[34]
Ziming Wang, Qianru Wu, Baolin Zheng, Junjie Wang, Kaiyu Huang, and Yanjie Shi. 2023. Sequence As Genes: An User Behavior Modeling Framework for Fraud Transaction Detection in E-commerce. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 5194–5203
2023
-
[35]
Xue Xia, Pong Eksombatchai, Nikil Pancha, Dhruvil Deven Badani, Po-Wei Wang, Neng Gu, Saurabh Vishwas Joshi, Nazanin Farahpour, Zhiyuan Zhang, and An- drew Zhai. 2023. TransAct: Transformer-based Realtime User Action Model for Recommendation at Pinterest. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 5249–5259
2023
-
[36]
Chen, Jimeng Sun, Jian Wu, and Jintai Chen
Jiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu, Danny Z. Chen, Jimeng Sun, Jian Wu, and Jintai Chen. 2024. Making Pre-trained Language Models Great on Tabular Prediction. arXiv:2403.01841 [cs.CL] https://arxiv.org/abs/2403.01841
Pith/arXiv arXiv 2024
-
[37]
Chin-Chia Michael Yeh, Uday Singh Saini, Xin Dai, Xiran Fan, Shubham Jain, Yujie Fan, Jiarui Sun, Junpeng Wang, Menghai Pan, Yingtong Dou, et al. 2025. TREASURE: A Transformer-Based Foundation Model for High-Volume Transac- tion Understanding.arXiv preprint arXiv:2511.19693(2025)
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.