REVIEW 2 major objections 5 minor 39 references
Warm-refit challengers off-path, promote them only after a fixed paired NLL edge over the live incumbent, and you get better crypto forecasts with far fewer model swaps.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 03:16 UTC pith:TMEWZHP4
load-bearing objection Solid ops paper: delayed-label paired gate on recursive serving trajectories, not another LOB architecture, with clean baselines and honest scope. the 2 major comments →
Train Often, Deploy Selectively: Forward-Gated Model Replacement in Crypto Markets
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Shadow Before Swap improves the recursive serving trajectory of probabilistic LOB forecasts relative to calendar replacement, schedule-matched automatic promotion, and continuous maintenance, while cutting deployed model changes by roughly four-fifths. On two nonoverlapping Binance episodes totaling 48 UTC weeks, three seeds, eight underlyings, and two perpetual contract types, SBS reduces hierarchically aggregated NLL by 0.1472%, 0.0755%, and 0.0428% against those three baselines, with positive episode-stratified four-week block intervals, by promoting 114 of 528 challengers.
What carries the argument
Shadow Before Swap (SBS): a causal shadow trial that deep-copies the full incumbent state, warm-fits the clone off the serving path, advances both branches on the same next week of delayed labels, and authorizes promotion only when the paired mean NLL advantage clears a fixed deadband (τ = 10^{-4}). The continuously maintained incumbent is the deployment null; a schedule-matched blind policy isolates the value of that authorization decision.
Load-bearing premise
A one-week paired NLL advantage under this fixed deadband and head-only delayed maintenance, measured in historical equal-weighted replay, is enough to decide which refits should become the parent of future live updates.
What would settle it
Replay the same recursive policies on a held-out nonoverlapping market episode (or live frozen-weight deployment) and check whether SBS’s NLL edge over blind promotion and continuous maintenance stays positive with the locked one-week trial and τ = 10^{-4}; a zero or negative maintenance contrast with many accepted bad promotions would falsify the claim.
If this is right
- Retraining cadence and release authorization can be separated: pipelines may propose refits on schedule while a forward gate controls which state enters service.
- Schedule-matched waiting alone is not enough; the paper’s blind comparator shows most of the remaining gain comes from refusing weak challengers.
- Accepted refits can beat a strong online-updated incumbent, so the safest policy is not simply never to full-refit.
- Deployed-state turnover can fall by about 78% without sacrificing probabilistic forecast quality under the reported setup.
- When candidate generators are badly misspecified, the same gate can keep serving near the incumbent by rejecting almost all proposals.
Where Pith is reading between the lines
- Any domain with delayed labels and a continuously adapting serving state—not only crypto LOBs—could reuse the same clone-shadow-compare release pattern between training and registry.
- Organizations that price rollback and validation work highly may want a wider deadband: the paper’s margin grid already shows fewer promotions with stable sign, which hints at a tunable ops cost knob.
- Equal asset and contract weighting makes this a governance estimand; capital-weighted live value would need weights frozen before outcomes and full recursive replay, as the author flags but does not run.
- The Temporal-CNN stress case suggests forward authorization is especially useful as a safety layer when validation repeatedly prefers unsafe replacements.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Shadow Before Swap (SBS), a causal deployment policy for production forecasters: at each scheduled boundary a challenger is deep-copied from the maintained incumbent, warm-refit off path, advanced with the incumbent on the same next week of delayed labels, and promoted only if paired mean NLL improves by at least τ=10^{-4}. Four recursive replays (maintenance, calendar, schedule-matched blind promotion, SBS) isolate timing, authorization, and the value of accepted refits. On two nonoverlapping Binance episodes (48 UTC weeks, 3 seeds, 8 underlyings, USD-M and COIN-M), SBS reduces hierarchically aggregated NLL by 0.1472% vs calendar, 0.0755% vs blind, and 0.0428% vs maintenance, with positive episode-stratified four-week block intervals, while accepting 114/528 challengers (78.4% fewer deployed changes). Directional consistency is reported across seeds, margins, trial lengths, a 20-asset panel, a supervised objective, and a Temporal-CNN failure-containment stress test.
Significance. If the reported recursive contrasts hold, the paper cleanly separates retraining from release authorization—an operational decision that production-ML and model-risk guidance often leave underspecified for delayed-label, nonstationary streams. Strengths include the schedule-matched blind contrast, full state isolation, development freeze of the gate before the primary episodes, multi-seed and multi-episode replication, margin/duration grids, block-bootstrap inference with episode stratification, and an explicit safety containment case under a misspecified Temporal-CNN. The contribution is a deployment policy rather than a new architecture, which is appropriately scoped and useful to applied forecasting systems. The work does not claim universal optimality or direct trading P&L; its value is a falsifiable, replayable authorization rule with measured turnover reduction.
major comments (2)
- [§3.5, Eq. (5), §5.3] §3.5 and Eq. (5): the primary estimand equal-weights underlyings, contract types, and weeks. That is a coherent release-policy estimand, but the practical claim in the abstract and §5.3 (a deployment policy that improves serving quality) is left somewhat open without at least one complete recursive sensitivity under frozen volume- or risk-based weights. A single pre-specified capital-weighted replay would show whether the sign and the 78.4% turnover reduction survive the weighting that many production desks would actually use; absence of that check is the main load-bearing gap between the reported NLL contrasts and operational adoption.
- [§4.1, Table 1, §5.2] §4.1 and §5.2: absolute gains vs maintenance are ~4.55×10^{-4} NLL per forecast. The conversion to ~455 log-loss units per million predictions and the matched-coverage Brier/diagnostic are helpful, but the manuscript still leans on relative percentages that are easy to over-read. Please state absolute NLL (or average per-week NLL) for all four policies in Table 1, and keep the service-scale interpretation explicitly tied to proper scoring rather than implying decision value without a frozen downstream rule—consistent with the limitation already noted in §5.4.
minor comments (5)
- [Figure 1, §2.1] Figure 1 pseudocode step (2) uses W', b', μ', σ' without defining the incumbent (W,b,μ,σ) in the main text notation of §2.1; a one-line alignment with S_t=(ϕ_t,W_t,N_t,Q_t,O_t) would help.
- [§3.1] §3.1: “development-selected and frozen” is carefully worded; still, briefly state whether any hyperparameter of the learner (Adam lr, patience, 28-day window) was touched after the lock, or only the gate rule.
- [Table 2, §4.1] Table 2: P/R counts are useful; adding the corresponding promotion rate next to the 48-week 114/528 figure in the main text would make the turnover claim easier to audit.
- [§6] §6 / Table 4: ChaCha and the Digalakis et al. challenger-switching framework are well placed; a sentence on whether progressive-validation bounds could replace the fixed τ deadband would round out the related-work contrast.
- [§2.1] Minor copy edits: “b𝑣_b” / “b𝐴_b” rendering in Eq. (2)–(3) is hard to read in the text build; “Forward–Forward” hyphenation is inconsistent in one place.
Circularity Check
No significant circularity: empirical policy contrasts on held-out recursive trajectories, not a self-referential derivation.
full rationale
SBS is defined as a causal shadow gate (clone, warm-refit, paired forward NLL, promote iff mean advantage ≥ τ) and evaluated by complete recursive serving trajectories against three distinct comparators—calendar replacement, schedule-matched blind promotion, and continuous maintenance—on external Binance market tapes. The SBS–blind contrast isolates authorization from waiting; SBS–maintenance tests whether accepted refits add value beyond never refitting; none of these estimands equals the gate rule by construction. τ=10^{-4} was development-selected and frozen before the primary episodes, with full recursive replays across a margin grid remaining positive, so the headline is not forced by the deadband. There are no load-bearing self-citations, uniqueness imports, or renamed known theorems. Standard empirical ML practice (frozen rule, OOT episodes, sensitivity) is not circularity under the stated criteria.
Axiom & Free-Parameter Ledger
free parameters (4)
- promotion deadband τ =
10^{-4}
- forward trial length =
1 week
- warm-refit and online head hyperparameters =
e.g. head SGD 3e-4; Adam 1e-3; normalizer rate 0.01; 24-channel net
- neutral return band and label delay =
5 bp; 300 s
axioms (5)
- domain assumption Prequential delayed-label evaluation: predict before consuming the current label; labels available only after fixed delay.
- domain assumption NLL (strictly proper) on equal asset/contract/week aggregation is the right estimand of deployment quality.
- domain assumption Deep-copied challenger/incumbent branches with separate normalizers, queues, and head updates correctly implement production champion–challenger isolation.
- ad hoc to paper Incumbent maintenance updates only the linear head via plain SGD while representation refreshes occur only on scheduled warm refits.
- standard math Circular moving-block bootstrap on aggregate weeks quantifies temporal robustness of realized recursive trajectories without resimulating counterfactual market orderings.
invented entities (1)
-
Shadow Before Swap (SBS) release policy
no independent evidence
read the original abstract
Production forecasting systems retrain models regularly, but a retrained candidate does not necessarily outperform a continuously maintained incumbent that has continued to learn. We introduce Shadow Before Swap (SBS), a deployment policy that warm-refits a challenger off the serving path, evaluates it against the maintained incumbent on the same next week of delayed labels, and promotes it only after a fixed paired negative-log-likelihood (NLL) advantage. In historical replay over two nonoverlapping Binance episodes spanning 48 UTC weeks, three seeds, eight underlyings, and two perpetual-futures contract types, SBS reduces NLL by 0.1472% relative to calendar replacement, 0.0755% relative to schedule-matched automatic promotion, and 0.0428% relative to continuous maintenance. The corresponding episode-stratified four-week block intervals are 0.1139%-0.1754%, 0.0521%-0.0980%, and 0.0301%-0.0554%, respectively. SBS promotes 114 of 528 challengers, reducing deployed model changes by 78.4% while improving the serving trajectory. The effect remains directionally consistent across seeds, trial budgets, promotion margins, an earlier 20-asset panel, and a topology-matched supervised objective. SBS thus provides a practical deployment policy that improves probabilistic forecasts while limiting consequential model-state transitions.
Figures
Reference graph
Works this paper leans on
-
[1]
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Xin Zhang, and Martin Zinkevich. 2017. TFX: A TensorFlow-...
2017
-
[2]
Binance. n.d. Binance Public Data. GitHub repository. Accessed July 19, 2026. https://github.com/binance/binance-public-data
2026
-
[3]
Board of Governors of the Federal Reserve System, Office of the Comptroller of the Currency, and Federal Deposit Insurance Corporation. 2026. Revised Guidance on Model Risk Management. Supervision and Regulation Letter SR 26-2. April 17, 2026. https://www.federalreserve.gov/supervisionreg/srletters/ SR2602.htm
2026
-
[4]
Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. 2017. The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction. In2017 IEEE International Conference on Big Data. IEEE, Piscataway, NJ, 1123–
2017
-
[5]
Glenn W. Brier. 1950. Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review78, 1 (1950), 1–3. doi:10.1175/1520-0493(1950)078<0001: VOFEIT>2.0.CO;2
-
[6]
Antonio Briola, Silvia Bartolucci, and Tomaso Aste. 2025. Deep Limit Order Book Forecasting: A Microstructural Guide.Quantitative Finance25, 7 (2025), 1101–1131. doi:10.1080/14697688.2025.2522911
arXiv 2025
-
[7]
C. K. Chow. 1970. On Optimum Recognition Error and Reject Tradeoff.IEEE Transactions on Information Theory16, 1 (1970), 41–46. doi:10.1109/TIT.1970. 1054406
doi:10.1109/tit.1970 1970
-
[8]
Botos Csaba, Wenxuan Zhang, Matthias Müller, Ser-Nam Lim, Mohamed Elho- seiny, Philip H. S. Torr, and Adel Bibi. 2024. Label Delay in Online Continual Learning. InAdvances in Neural Information Processing Systems, Vol. 37. Curran Associates, Inc., Vancouver, Canada, 119976–120012. doi:10.52202/079017-3813
-
[9]
Zhenwen Dai, Praveen Chandar, Ghazal Fazelnia, Benjamin Carterette, and Mounia Lalmas. 2020. Model Selection for Production System via Automated Online Experiments. InAdvances in Neural Information Processing Systems, Vol. 33. Curran Associates, Inc., Virtual Event, 1106–1116. https://proceedings.neurips. cc/paper/2020/hash/0c72cb7ee1512f800abe27823a792d0...
2020
-
[10]
Štěpán Davidovič and Betsy Beyer. 2018. Canary Analysis Service: Automated Ca- narying Quickens Development, Improves Production Safety, and Helps Prevent Outages.ACM Queue16, 1 (2018), 35–57. doi:10.1145/3194653.3194655
arXiv 2018
-
[11]
Philip Dawid
A. Philip Dawid. 1984. Present Position and Potential Developments: Some Personal Views: Statistical Theory: The Prequential Approach.Journal of the Royal Statistical Society: Series A (General)147, 2 (1984), 278–290. doi:10.2307/ 2981683
1984
-
[12]
Francis X. Diebold and Roberto S. Mariano. 1995. Comparing Predictive Accuracy. Journal of Business & Economic Statistics13, 3 (1995), 253–263. doi:10.1080/ 07350015.1995.10524599
arXiv 1995
-
[13]
Vassilis Digalakis Jr, Christophe Pérignon, Sébastien Saurin, and Flore Sentenac
-
[14]
Ran El-Yaniv and Yair Wiener. 2010. On the Foundations of Noise-Free Selective Classification.Journal of Machine Learning Research11, 53 (2010), 1605–1641. https://jmlr.org/papers/v11/el-yaniv10a.html
2010
-
[15]
Yonatan Geifman and Ran El-Yaniv. 2017. Selective Classification for Deep Neural Networks. InAdvances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc., Long Beach, California, USA, 4878–4887
2017
-
[16]
Raffaella Giacomini and Halbert White. 2006. Tests of Conditional Predictive Ability.Econometrica74, 6 (2006), 1545–1578. doi:10.1111/j.1468-0262.2006.00718. x
arXiv 2006
-
[17]
Tilmann Gneiting and Adrian E. Raftery. 2007. Strictly Proper Scoring Rules, Prediction, and Estimation.J. Amer. Statist. Assoc.102, 477 (2007), 359–378. doi:10.1198/016214506000001437
-
[18]
Geoffrey Hinton. 2022. The Forward-Forward Algorithm: Some Preliminary Investigations. arXiv:2212.13345 [cs.LG] https://arxiv.org/abs/2212.13345
Pith/arXiv arXiv 2022
-
[19]
Pooria Joulani, Andras Gyorgy, and Csaba Szepesvari. 2013. Online Learning under Delayed Feedback. InProceedings of the 30th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 28), Sanjoy Dasgupta and David McAllester (Eds.). PMLR, Atlanta, Georgia, USA, 1453–1461. Issue 3. https://proceedings.mlr.press/v28/joulani13.html
2013
-
[20]
S. N. Lahiri. 2003.Resampling Methods for Dependent Data. Springer, New York. doi:10.1007/978-1-4757-3803-2
-
[21]
Whitney K. Newey and Kenneth D. West. 1987. A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix.Econo- metrica55, 3 (1987), 703–708. doi:10.2307/1913610
doi:10.2307/1913610 1987
-
[22]
Politis and Joseph P
Dimitris N. Politis and Joseph P. Romano. 1992. A Circular Block-Resampling Procedure for Stationary Data. InExploring the Limits of Bootstrap, Raoul LePage and Lynne Billard (Eds.). John Wiley & Sons, New York, 263–270
1992
-
[23]
Dimitris N. Politis and Joseph P. Romano. 1994. The Stationary Bootstrap.J. Amer. Statist. Assoc.89, 428 (1994), 1303–1313. doi:10.1080/01621459.1994.10476870
arXiv 1994
-
[24]
Matteo Prata, Giuseppe Masi, Leonardo Berti, Viviana Arrigoni, Andrea Coletta, Irene Cannistraci, Svitlana Vyetrenko, Paola Velardi, and Novella Bartolini. 2024. LOB-Based Deep Learning Models for Stock Price Trend Prediction: A Benchmark Study.Artificial Intelligence Review57, 5, Article 116 (2024), 45 pages. doi:10. 1007/s10462-024-10715-4
2024
-
[25]
Florence Regol, Leo Schwinn, Kyle Sprague, Mark Coates, and Thomas Markovich
-
[26]
Mohammad Reza Karimi, Nezihe Merve Gürel, Bojan Karlaš, Johannes Rausch, Ce Zhang, and Andreas Krause. 2021. Online Active Model Selection for Pre-trained Classifiers. InProceedings of the 24th International Conference on Artificial Intelli- gence and Statistics (Proceedings of Machine Learning Research, Vol. 130). PMLR, Virtual Event, 307–315. https://pr...
2021
-
[27]
Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, and Matthias Bethge. 2020. Improving Robustness Against Com- mon Corruptions by Covariate Shift Adaptation. InAdvances in Neural In- formation Processing Systems, Vol. 33. Curran Associates, Inc., Virtual Event, 11539–11551. https://proceedings.neurips.cc/paper_files/paper/202...
2020
-
[28]
Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Denni- son
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Denni- son. 2015. Hidden Technical Debt in Machine Learning Systems. InAdvances in Neural Information Processing Systems, Vol. 28. Curran Associates, Inc., Red Hook, NY, 2503–2511. https://papers.nips.cc/paper/...
2015
-
[29]
Efros, and Moritz Hardt
Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, and Moritz Hardt. 2020. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts. InProceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119). PMLR, Virtual Event, 9229–9248. https://proceedings.mlr....
2020
-
[30]
InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol
When to Retrain a Machine Learning Model. InProceedings of the 42nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267). PMLR, Vancouver, Canada, 51369–51404. https://proceedings. mlr.press/v267/regol25a.html
-
[31]
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. 2021. Tent: Fully Test-Time Adaptation by Entropy Minimization. In International Conference on Learning Representations. OpenReview.net, Virtual Event, 15 pages. https://openreview.net/forum?id=uXl3bZLkr3c
2021
-
[32]
Qingyun Wu, Chi Wang, John Langford, Paul Mineiro, and Marco Rossi. 2021. ChaCha for Online AutoML. InProceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139). PMLR, Virtual Event, 11263–11273. https://proceedings.mlr.press/v139/wu21d.html
2021
-
[33]
Yue Xiao, Carmine Ventre, Yuhan Wang, Haochen Li, Yuxi Huan, and Buhong Liu. 2025. LiT: Limit Order Book Transformer.Frontiers in Artificial Intelligence 8, Article 1616485 (2025), 12 pages. doi:10.3389/frai.2025.1616485 Train Often, Deploy Selectively: Forward-Gated Model Replacement in Crypto Markets
arXiv 2025
-
[34]
Zihao Zhang, Stefan Zohren, and Stephen Roberts. 2019. DeepLOB: Deep Con- volutional Neural Networks for Limit Order Books.IEEE Transactions on Signal Processing67, 11 (2019), 3001–3012. doi:10.1109/TSP.2019.2907260
arXiv 2019
-
[35]
2023.Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Elham Tabassi. 2023.Artificial Intelligence Risk Management Framework (AI RMF 1.0). Technical Report NIST AI 100-1. National Institute of Standards and Technology. doi:10.6028/NIST.AI.100-1
-
[1132]
doi:10.1109/BigData.2017.8258038
arXiv 2017
-
[1395]
doi:10.1145/3097983.3098021
-
[2025]
FIN-2025-1601
The Challenger: When Do New Data Sources Justify Switching Machine Learning Models? HEC Paris Research Paper No. FIN-2025-1601. Revised July 1,
2025
-
[2026]
doi:10.2139/ssrn.5946174
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.