Pith. sign in

REVIEW 2 major objections 5 minor 39 references

Warm-refit challengers off-path, promote them only after a fixed paired NLL edge over the live incumbent, and you get better crypto forecasts with far fewer model swaps.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 03:16 UTC pith:TMEWZHP4

load-bearing objection Solid ops paper: delayed-label paired gate on recursive serving trajectories, not another LOB architecture, with clean baselines and honest scope. the 2 major comments →

arxiv 2607.28577 v1 pith:TMEWZHP4 submitted 2026-07-30 cs.CE q-fin.TR

Train Often, Deploy Selectively: Forward-Gated Model Replacement in Crypto Markets

classification cs.CE q-fin.TR
keywords model replacementnonstationarityshadow evaluationdelayed labelslimit order booksprobabilistic forecastingonline learningdeployment policy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Production forecasters keep learning between rebuilds, so a newly trained model can beat an old checkpoint and still lose to the state that is actually serving. This paper argues that replacement is therefore a sequential authorization problem, not just a training schedule. It proposes Shadow Before Swap: clone the incumbent, warm-refit the clone on mature history, run challenger and incumbent side by side on the next week of delayed labels, and promote only if the challenger’s mean negative log-likelihood is better by a fixed deadband. In 48 weeks of Binance perpetual-futures replay across seeds, underlyings, and contract types, that gate beats calendar replacement, blind promotion on the same schedule, and pure continuous maintenance, while accepting only about one in five challengers. A sympathetic reader cares because the policy sits between an existing retrain pipeline and the registry: you can train often, yet install only when forward evidence justifies changing the parent of every later update.

Core claim

Shadow Before Swap improves the recursive serving trajectory of probabilistic LOB forecasts relative to calendar replacement, schedule-matched automatic promotion, and continuous maintenance, while cutting deployed model changes by roughly four-fifths. On two nonoverlapping Binance episodes totaling 48 UTC weeks, three seeds, eight underlyings, and two perpetual contract types, SBS reduces hierarchically aggregated NLL by 0.1472%, 0.0755%, and 0.0428% against those three baselines, with positive episode-stratified four-week block intervals, by promoting 114 of 528 challengers.

What carries the argument

Shadow Before Swap (SBS): a causal shadow trial that deep-copies the full incumbent state, warm-fits the clone off the serving path, advances both branches on the same next week of delayed labels, and authorizes promotion only when the paired mean NLL advantage clears a fixed deadband (τ = 10^{-4}). The continuously maintained incumbent is the deployment null; a schedule-matched blind policy isolates the value of that authorization decision.

Load-bearing premise

A one-week paired NLL advantage under this fixed deadband and head-only delayed maintenance, measured in historical equal-weighted replay, is enough to decide which refits should become the parent of future live updates.

What would settle it

Replay the same recursive policies on a held-out nonoverlapping market episode (or live frozen-weight deployment) and check whether SBS’s NLL edge over blind promotion and continuous maintenance stays positive with the locked one-week trial and τ = 10^{-4}; a zero or negative maintenance contrast with many accepted bad promotions would falsify the claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Retraining cadence and release authorization can be separated: pipelines may propose refits on schedule while a forward gate controls which state enters service.
  • Schedule-matched waiting alone is not enough; the paper’s blind comparator shows most of the remaining gain comes from refusing weak challengers.
  • Accepted refits can beat a strong online-updated incumbent, so the safest policy is not simply never to full-refit.
  • Deployed-state turnover can fall by about 78% without sacrificing probabilistic forecast quality under the reported setup.
  • When candidate generators are badly misspecified, the same gate can keep serving near the incumbent by rejecting almost all proposals.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Any domain with delayed labels and a continuously adapting serving state—not only crypto LOBs—could reuse the same clone-shadow-compare release pattern between training and registry.
  • Organizations that price rollback and validation work highly may want a wider deadband: the paper’s margin grid already shows fewer promotions with stable sign, which hints at a tunable ops cost knob.
  • Equal asset and contract weighting makes this a governance estimand; capital-weighted live value would need weights frozen before outcomes and full recursive replay, as the author flags but does not run.
  • The Temporal-CNN stress case suggests forward authorization is especially useful as a safety layer when validation repeatedly prefers unsafe replacements.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes Shadow Before Swap (SBS), a causal deployment policy for production forecasters: at each scheduled boundary a challenger is deep-copied from the maintained incumbent, warm-refit off path, advanced with the incumbent on the same next week of delayed labels, and promoted only if paired mean NLL improves by at least τ=10^{-4}. Four recursive replays (maintenance, calendar, schedule-matched blind promotion, SBS) isolate timing, authorization, and the value of accepted refits. On two nonoverlapping Binance episodes (48 UTC weeks, 3 seeds, 8 underlyings, USD-M and COIN-M), SBS reduces hierarchically aggregated NLL by 0.1472% vs calendar, 0.0755% vs blind, and 0.0428% vs maintenance, with positive episode-stratified four-week block intervals, while accepting 114/528 challengers (78.4% fewer deployed changes). Directional consistency is reported across seeds, margins, trial lengths, a 20-asset panel, a supervised objective, and a Temporal-CNN failure-containment stress test.

Significance. If the reported recursive contrasts hold, the paper cleanly separates retraining from release authorization—an operational decision that production-ML and model-risk guidance often leave underspecified for delayed-label, nonstationary streams. Strengths include the schedule-matched blind contrast, full state isolation, development freeze of the gate before the primary episodes, multi-seed and multi-episode replication, margin/duration grids, block-bootstrap inference with episode stratification, and an explicit safety containment case under a misspecified Temporal-CNN. The contribution is a deployment policy rather than a new architecture, which is appropriately scoped and useful to applied forecasting systems. The work does not claim universal optimality or direct trading P&L; its value is a falsifiable, replayable authorization rule with measured turnover reduction.

major comments (2)
  1. [§3.5, Eq. (5), §5.3] §3.5 and Eq. (5): the primary estimand equal-weights underlyings, contract types, and weeks. That is a coherent release-policy estimand, but the practical claim in the abstract and §5.3 (a deployment policy that improves serving quality) is left somewhat open without at least one complete recursive sensitivity under frozen volume- or risk-based weights. A single pre-specified capital-weighted replay would show whether the sign and the 78.4% turnover reduction survive the weighting that many production desks would actually use; absence of that check is the main load-bearing gap between the reported NLL contrasts and operational adoption.
  2. [§4.1, Table 1, §5.2] §4.1 and §5.2: absolute gains vs maintenance are ~4.55×10^{-4} NLL per forecast. The conversion to ~455 log-loss units per million predictions and the matched-coverage Brier/diagnostic are helpful, but the manuscript still leans on relative percentages that are easy to over-read. Please state absolute NLL (or average per-week NLL) for all four policies in Table 1, and keep the service-scale interpretation explicitly tied to proper scoring rather than implying decision value without a frozen downstream rule—consistent with the limitation already noted in §5.4.
minor comments (5)
  1. [Figure 1, §2.1] Figure 1 pseudocode step (2) uses W', b', μ', σ' without defining the incumbent (W,b,μ,σ) in the main text notation of §2.1; a one-line alignment with S_t=(ϕ_t,W_t,N_t,Q_t,O_t) would help.
  2. [§3.1] §3.1: “development-selected and frozen” is carefully worded; still, briefly state whether any hyperparameter of the learner (Adam lr, patience, 28-day window) was touched after the lock, or only the gate rule.
  3. [Table 2, §4.1] Table 2: P/R counts are useful; adding the corresponding promotion rate next to the 48-week 114/528 figure in the main text would make the turnover claim easier to audit.
  4. [§6] §6 / Table 4: ChaCha and the Digalakis et al. challenger-switching framework are well placed; a sentence on whether progressive-validation bounds could replace the fixed τ deadband would round out the related-work contrast.
  5. [§2.1] Minor copy edits: “b𝑣_b” / “b𝐴_b” rendering in Eq. (2)–(3) is hard to read in the text build; “Forward–Forward” hyphenation is inconsistent in one place.

Circularity Check

0 steps flagged

No significant circularity: empirical policy contrasts on held-out recursive trajectories, not a self-referential derivation.

full rationale

SBS is defined as a causal shadow gate (clone, warm-refit, paired forward NLL, promote iff mean advantage ≥ τ) and evaluated by complete recursive serving trajectories against three distinct comparators—calendar replacement, schedule-matched blind promotion, and continuous maintenance—on external Binance market tapes. The SBS–blind contrast isolates authorization from waiting; SBS–maintenance tests whether accepted refits add value beyond never refitting; none of these estimands equals the gate rule by construction. τ=10^{-4} was development-selected and frozen before the primary episodes, with full recursive replays across a margin grid remaining positive, so the headline is not forced by the deadband. There are no load-bearing self-citations, uniqueness imports, or renamed known theorems. Standard empirical ML practice (frozen rule, OOT episodes, sensitivity) is not circularity under the stated criteria.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 1 invented entities

Load-bearing content is mostly operational and empirical: a hand-chosen promotion deadband and trial budget, a specific online maintenance operator, delayed-label timing, NLL as the deployment objective, and equal decision-budget aggregation. No new physical entities. The scientific add is the gated recursive policy and replay evidence, not a parameter-free theory.

free parameters (4)
  • promotion deadband τ = 10^{-4}
    Development-selected operational threshold requiring mean paired NLL advantage before promote; primary results use this value though sensitivity grid is reported.
  • forward trial length = 1 week
    Primary evidence budget before the release decision; alternatives (2–3 weeks) are secondary and not promoted to a new rule.
  • warm-refit and online head hyperparameters = e.g. head SGD 3e-4; Adam 1e-3; normalizer rate 0.01; 24-channel net
    Adam/SGD rates, batch size, patience, normalizer rate, 28-day refit window, architecture widths—all chosen settings that define challenger/incumbent behavior and thus measured gaps.
  • neutral return band and label delay = 5 bp; 300 s
    Target definition (5 bp band, 300s horizon/delay) shapes NLL and gate inputs; fixed rather than derived.
axioms (5)
  • domain assumption Prequential delayed-label evaluation: predict before consuming the current label; labels available only after fixed delay.
    Eq. 1 and §2.1–2.2 ground causal timing; standard in delayed online learning but essential to the gate’s validity.
  • domain assumption NLL (strictly proper) on equal asset/contract/week aggregation is the right estimand of deployment quality.
    §3.5 and §5.2; monetary utility, risk weights, and trading frictions are explicitly not identified.
  • domain assumption Deep-copied challenger/incumbent branches with separate normalizers, queues, and head updates correctly implement production champion–challenger isolation.
    §2.2; required for interpreting policy contrasts as release effects rather than shared-state artifacts.
  • ad hoc to paper Incumbent maintenance updates only the linear head via plain SGD while representation refreshes occur only on scheduled warm refits.
    §3.3 defines the strong null; results are conditional on this maintenance mechanism, not arbitrary online learners.
  • standard math Circular moving-block bootstrap on aggregate weeks quantifies temporal robustness of realized recursive trajectories without resimulating counterfactual market orderings.
    §5.1 cites block bootstrap practice and correctly limits interpretation versus full counterfactual-history uncertainty.
invented entities (1)
  • Shadow Before Swap (SBS) release policy no independent evidence
    purpose: Authorize model-state replacement only after off-path paired forward NLL evidence against a maintained incumbent.
    Core proposed control rule (Eqs. 2–3, Fig. 1). It is an algorithmic policy, not a latent physical object; independent evidence is the paper’s own replays, not an external measurement channel.

pith-pipeline@v1.2.0-daily-grok45 · 18352 in / 3593 out tokens · 88903 ms · 2026-07-31T03:16:30.094724+00:00 · methodology

0 comments
read the original abstract

Production forecasting systems retrain models regularly, but a retrained candidate does not necessarily outperform a continuously maintained incumbent that has continued to learn. We introduce Shadow Before Swap (SBS), a deployment policy that warm-refits a challenger off the serving path, evaluates it against the maintained incumbent on the same next week of delayed labels, and promotes it only after a fixed paired negative-log-likelihood (NLL) advantage. In historical replay over two nonoverlapping Binance episodes spanning 48 UTC weeks, three seeds, eight underlyings, and two perpetual-futures contract types, SBS reduces NLL by 0.1472% relative to calendar replacement, 0.0755% relative to schedule-matched automatic promotion, and 0.0428% relative to continuous maintenance. The corresponding episode-stratified four-week block intervals are 0.1139%-0.1754%, 0.0521%-0.0980%, and 0.0301%-0.0554%, respectively. SBS promotes 114 of 528 challengers, reducing deployed model changes by 78.4% while improving the serving trajectory. The effect remains directionally consistent across seeds, trial budgets, promotion margins, an earlier 20-asset panel, and a topology-matched supervised objective. SBS thus provides a practical deployment policy that improves probabilistic forecasts while limiting consequential model-state transitions.

Figures

Figures reproduced from arXiv: 2607.28577 by Aditya Dutta.

Figure 1
Figure 1. Figure 1: SBS decision flow and pseudocode. The challenger [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Selective replacement replicates. (a) Episode and stratified relative effects; green annotations give exact pooled absolute [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: What the forward trial measures. (a) Binned density of one-week paired trial gain and three-week future challenger [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Breadth evidence. Primary time breadth, earlier [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Temporal-CNN failure containment. (a) Bootstrap-median SBS relative NLL reduction with pointwise 95% four-week [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 2 canonical work pages

  1. [1]

    Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Xin Zhang, and Martin Zinkevich. 2017. TFX: A TensorFlow-...

  2. [2]

    Binance. n.d. Binance Public Data. GitHub repository. Accessed July 19, 2026. https://github.com/binance/binance-public-data

  3. [3]

    Board of Governors of the Federal Reserve System, Office of the Comptroller of the Currency, and Federal Deposit Insurance Corporation. 2026. Revised Guidance on Model Risk Management. Supervision and Regulation Letter SR 26-2. April 17, 2026. https://www.federalreserve.gov/supervisionreg/srletters/ SR2602.htm

  4. [4]

    Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. 2017. The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. In2017 IEEE International Conference on Big Data. IEEE, Piscataway, NJ, 1123–

  5. [5]

    Glenn W. Brier. 1950. Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review78, 1 (1950), 1–3. doi:10.1175/1520-0493(1950)078<0001: VOFEIT>2.0.CO;2

  6. [6]

    Antonio Briola, Silvia Bartolucci, and Tomaso Aste. 2025. Deep Limit Order Book Forecasting: A Microstructural Guide.Quantitative Finance25, 7 (2025), 1101–1131. doi:10.1080/14697688.2025.2522911

  7. [7]

    C. K. Chow. 1970. On Optimum Recognition Error and Reject Tradeoff.IEEE Transactions on Information Theory16, 1 (1970), 41–46. doi:10.1109/TIT.1970. 1054406

  8. [8]

    Botos Csaba, Wenxuan Zhang, Matthias Müller, Ser-Nam Lim, Mohamed Elho- seiny, Philip H. S. Torr, and Adel Bibi. 2024. Label Delay in Online Continual Learning. InAdvances in Neural Information Processing Systems, Vol. 37. Curran Associates, Inc., Vancouver, Canada, 119976–120012. doi:10.52202/079017-3813

  9. [9]

    Zhenwen Dai, Praveen Chandar, Ghazal Fazelnia, Benjamin Carterette, and Mounia Lalmas. 2020. Model Selection for Production System via Automated Online Experiments. InAdvances in Neural Information Processing Systems, Vol. 33. Curran Associates, Inc., Virtual Event, 1106–1116. https://proceedings.neurips. cc/paper/2020/hash/0c72cb7ee1512f800abe27823a792d0...

  10. [10]

    Štěpán Davidovič and Betsy Beyer. 2018. Canary Analysis Service: Automated Ca- narying Quickens Development, Improves Production Safety, and Helps Prevent Outages.ACM Queue16, 1 (2018), 35–57. doi:10.1145/3194653.3194655

  11. [11]

    Philip Dawid

    A. Philip Dawid. 1984. Present Position and Potential Developments: Some Personal Views: Statistical Theory: The Prequential Approach.Journal of the Royal Statistical Society: Series A (General)147, 2 (1984), 278–290. doi:10.2307/ 2981683

  12. [12]

    Diebold and Roberto S

    Francis X. Diebold and Roberto S. Mariano. 1995. Comparing Predictive Accuracy. Journal of Business & Economic Statistics13, 3 (1995), 253–263. doi:10.1080/ 07350015.1995.10524599

  13. [13]

    Vassilis Digalakis Jr, Christophe Pérignon, Sébastien Saurin, and Flore Sentenac

  14. [14]

    Ran El-Yaniv and Yair Wiener. 2010. On the Foundations of Noise-Free Selective Classification.Journal of Machine Learning Research11, 53 (2010), 1605–1641. https://jmlr.org/papers/v11/el-yaniv10a.html

  15. [15]

    Yonatan Geifman and Ran El-Yaniv. 2017. Selective Classification for Deep Neural Networks. InAdvances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc., Long Beach, California, USA, 4878–4887

  16. [16]

    Raffaella Giacomini and Halbert White. 2006. Tests of Conditional Predictive Ability.Econometrica74, 6 (2006), 1545–1578. doi:10.1111/j.1468-0262.2006.00718. x

  17. [17]

    Tilmann Gneiting and Adrian E. Raftery. 2007. Strictly Proper Scoring Rules, Prediction, and Estimation.J. Amer. Statist. Assoc.102, 477 (2007), 359–378. doi:10.1198/016214506000001437

  18. [18]

    Geoffrey Hinton. 2022. The Forward-Forward Algorithm: Some Preliminary Investigations. arXiv:2212.13345 [cs.LG] https://arxiv.org/abs/2212.13345

  19. [19]

    Pooria Joulani, Andras Gyorgy, and Csaba Szepesvari. 2013. Online Learning under Delayed Feedback. InProceedings of the 30th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 28), Sanjoy Dasgupta and David McAllester (Eds.). PMLR, Atlanta, Georgia, USA, 1453–1461. Issue 3. https://proceedings.mlr.press/v28/joulani13.html

  20. [20]

    S. N. Lahiri. 2003.Resampling Methods for Dependent Data. Springer, New York. doi:10.1007/978-1-4757-3803-2

  21. [21]

    Newey and Kenneth D

    Whitney K. Newey and Kenneth D. West. 1987. A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix.Econo- metrica55, 3 (1987), 703–708. doi:10.2307/1913610

  22. [22]

    Politis and Joseph P

    Dimitris N. Politis and Joseph P. Romano. 1992. A Circular Block-Resampling Procedure for Stationary Data. InExploring the Limits of Bootstrap, Raoul LePage and Lynne Billard (Eds.). John Wiley & Sons, New York, 263–270

  23. [23]

    Politis and Joseph P

    Dimitris N. Politis and Joseph P. Romano. 1994. The Stationary Bootstrap.J. Amer. Statist. Assoc.89, 428 (1994), 1303–1313. doi:10.1080/01621459.1994.10476870

  24. [24]

    Matteo Prata, Giuseppe Masi, Leonardo Berti, Viviana Arrigoni, Andrea Coletta, Irene Cannistraci, Svitlana Vyetrenko, Paola Velardi, and Novella Bartolini. 2024. LOB-Based Deep Learning Models for Stock Price Trend Prediction: A Benchmark Study.Artificial Intelligence Review57, 5, Article 116 (2024), 45 pages. doi:10. 1007/s10462-024-10715-4

  25. [25]

    Florence Regol, Leo Schwinn, Kyle Sprague, Mark Coates, and Thomas Markovich

  26. [26]

    Mohammad Reza Karimi, Nezihe Merve Gürel, Bojan Karlaš, Johannes Rausch, Ce Zhang, and Andreas Krause. 2021. Online Active Model Selection for Pre-trained Classifiers. InProceedings of the 24th International Conference on Artificial Intelli- gence and Statistics (Proceedings of Machine Learning Research, Vol. 130). PMLR, Virtual Event, 307–315. https://pr...

  27. [27]

    Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, and Matthias Bethge. 2020. Improving Robustness Against Com- mon Corruptions by Covariate Shift Adaptation. InAdvances in Neural In- formation Processing Systems, Vol. 33. Curran Associates, Inc., Virtual Event, 11539–11551. https://proceedings.neurips.cc/paper_files/paper/202...

  28. [28]

    Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Denni- son

    D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Denni- son. 2015. Hidden Technical Debt in Machine Learning Systems. InAdvances in Neural Information Processing Systems, Vol. 28. Curran Associates, Inc., Red Hook, NY, 2503–2511. https://papers.nips.cc/paper/...

  29. [29]

    Efros, and Moritz Hardt

    Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, and Moritz Hardt. 2020. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts. InProceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119). PMLR, Virtual Event, 9229–9248. https://proceedings.mlr....

  30. [30]

    InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol

    When to Retrain a Machine Learning Model. InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267). PMLR, Vancouver, Canada, 51369–51404. https://proceedings. mlr.press/v267/regol25a.html

  31. [31]

    Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. 2021. Tent: Fully Test-Time Adaptation by Entropy Minimization. In International Conference on Learning Representations. OpenReview.net, Virtual Event, 15 pages. https://openreview.net/forum?id=uXl3bZLkr3c

  32. [32]

    Qingyun Wu, Chi Wang, John Langford, Paul Mineiro, and Marco Rossi. 2021. ChaCha for Online AutoML. InProceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139). PMLR, Virtual Event, 11263–11273. https://proceedings.mlr.press/v139/wu21d.html

  33. [33]

    Yue Xiao, Carmine Ventre, Yuhan Wang, Haochen Li, Yuxi Huan, and Buhong Liu. 2025. LiT: Limit Order Book Transformer.Frontiers in Artificial Intelligence 8, Article 1616485 (2025), 12 pages. doi:10.3389/frai.2025.1616485 Train Often, Deploy Selectively: Forward-Gated Model Replacement in Crypto Markets

  34. [34]

    Zihao Zhang, Stefan Zohren, and Stephen Roberts. 2019. DeepLOB: Deep Con- volutional Neural Networks for Limit Order Books.IEEE Transactions on Signal Processing67, 11 (2019), 3001–3012. doi:10.1109/TSP.2019.2907260

  35. [35]

    2023.Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    Elham Tabassi. 2023.Artificial Intelligence Risk Management Framework (AI RMF 1.0). Technical Report NIST AI 100-1. National Institute of Standards and Technology. doi:10.6028/NIST.AI.100-1

  36. [1132]

    doi:10.1109/BigData.2017.8258038

  37. [1395]

    doi:10.1145/3097983.3098021

  38. [2025]

    FIN-2025-1601

    The Challenger: When Do New Data Sources Justify Switching Machine Learning Models? HEC Paris Research Paper No. FIN-2025-1601. Revised July 1,

  39. [2026]

    doi:10.2139/ssrn.5946174