{"id":"7f5b5e3f-a5d2-4683-ac68-3f0ffab2b2fb","arxiv_id":"2607.00475","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"End-to-end AI policies for cross-asset futures timing outperform rules-based benchmarks on pooled portfolios but vary by asset class, with transformers showing better cost-adjusted performance than LSTMs.","lead":"The paper trains LSTM and transformer models to map market states directly to weights across 16 CME futures using a differentiable Sharpe ratio loss, finding the learned policies beat equal weighting, risk parity, and time-series momentum on the pooled portfolio and some sub-classes but not uniformly, with the transformer handling costs better by trading less. A smart generalist might read it to see whether end-to-end AI adds value over simple rules in cross-asset timing when","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Out-of-sample ranking advantage may be an artifact of untested training window, features, or hyperparameters","rationale":"The reader's weakest_assumption exactly isolates the load-bearing empirical-finance risk. No stronger internal inconsistency appears from the abstract alone, and the full-text placeholder does not alter the fact that robustness to data choices remains unaddressed.","tokens_in":1729,"tokens_out":300,"duration_ms":21355,"concrete_test":"Re-train both architectures on the same 16 futures but with (a) training window ending 24 months earlier and (b) an alternative feature set that drops all volatility and volume inputs; recompute the pooled Sharpe ranking versus the three rules. If the learned-policy advantage reverses or becomes statistically insignificant in either case, the headline result is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the LSTM/transformer policies' pooled and sub-asset outperformance (gross and net of costs) over equal-weight, risk-parity, and TSMOM is not driven by the specific CME futures sample period, the exact market-state features, or hyperparameter/search choices. Because futures returns are regime-dependent and non-stationary, any of these can produce spurious ranking advantages that vanish under modest perturbations. The abstract supplies no information on walk-forward validation, multiple random seeds, or ablation on feature sets, so the reported edge cannot be distinguished from an in-sample artifact.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes training LSTM and transformer models as end-to-end parametric policies that map market-state features directly to weights across 16 liquid CME futures contracts, using a differentiable Sharpe-ratio objective. These policies are benchmarked against equal-weight, risk-parity, and time-series momentum rules; the abstract reports that the learned policies rank above the rules on the pooled cross-asset portfolio and in several sub-asset classes, with the transformer outperforming the LSTM once transaction costs are included because of substantially lower turnover.","tokens_in":1851,"tokens_out":401,"duration_ms":26782,"significance":"If the reported ranking advantages survive proper out-of-sample validation, the work would provide concrete evidence that direct policy learning can improve upon standard rule-based timing strategies in futures markets and would highlight the practical importance of turnover when selecting among neural architectures. The differentiable-Sharpe training approach is a methodological strength that enables truly end-to-end optimization.","major_comments":[{"comment":"Abstract: the central empirical claim—that the LSTM and transformer policies produce higher pooled and sub-asset rankings than equal-weight, risk-parity, and TSMOM, both gross and net of costs—cannot be evaluated because the abstract (and the supplied text) supplies no information on the sample period, walk-forward or cross-validation scheme, number of random seeds, or any statistical significance tests for the reported rankings.","section":"Abstract"},{"comment":"The claim that the transformer’s lower turnover produces a net-of-cost advantage that “matches or exceeds equal weighting” is load-bearing for the paper’s conclusion about when AI models beat rules; without reported ablation on feature sets, hyperparameter sensitivity, or multiple training windows, it is impossible to rule out that the advantage is an artifact of the particular CME futures sample or market-state construction.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful review and for identifying points where additional detail will strengthen the manuscript. We respond to each major comment below.","responses":[{"response":"We agree that the abstract and main text must supply these details for the empirical claims to be evaluable. We will revise the abstract to summarize the sample period, the walk-forward validation procedure, the number of random seeds, and the statistical tests performed on the rankings. We will also ensure the methodology section states these elements explicitly and prominently.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central empirical claim—that the LSTM and transformer policies produce higher pooled and sub-asset rankings than equal-weight, risk-parity, and TSMOM, both gross and net of costs—cannot be evaluated because the abstract (and the supplied text) supplies no information on the sample period, walk-forward or cross-validation scheme, number of random seeds, or any statistical significance tests for the reported rankings."},{"response":"We agree that additional robustness checks are warranted to support the load-bearing claim. The manuscript already reports turnover and net-of-cost metrics across asset classes, but we will add an appendix containing ablations on feature sets, hyperparameter grids, and results from alternative training windows. These additions will allow readers to assess whether the transformer advantage is robust or sample-specific.","revision_made":"yes","referee_comment":"[Abstract] The claim that the transformer’s lower turnover produces a net-of-cost advantage that “matches or exceeds equal weighting” is load-bearing for the paper’s conclusion about when AI models beat rules; without reported ablation on feature sets, hyperparameter sensitivity, or multiple training windows, it is impossible to rule out that the advantage is an artifact of the particular CME futures sample or market-state construction."}],"tokens_in":1351,"tokens_out":398,"duration_ms":37513,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper applies end-to-end differentiable policy learning to timing across liquid futures and reports that the transformer version holds an edge over equal-weight, risk-parity, and TSMOM once transaction costs enter, mainly because it trades less than the LSTM.\n\nWhat is new is the direct architecture comparison in this narrow setting, with explicit attention to turnover rather than just gross Sharpe. The work does a decent job keeping the focus on practical, liquid instruments where any edge has to survive real frictions.\n\nThe soft spot is the total absence of information on data periods, walk-forward splits, statistical tests, or hyperparameter robustness. Futures returns are regime-dependent, so without those checks the reported ranking advantage could easily be tied to the particular window or feature choices rather than a stable improvement. The stress-test concern lands because the abstract gives nothing to distinguish signal from artifact.\n\nThis paper is for quants who already run timing models on futures and want a data point on whether moving from rules to learned policies changes the cost picture. A reader already working in differentiable optimization or cross-asset allocation would get the most out of it.\n\nIt deserves a serious referee. The question is concrete and the domain is relevant; if the full methods section shows proper out-of-sample protocols and the results survive modest perturbations, the cost comparison would be worth having in the literature.\n\nI would send it to peer review so the authors can supply the missing validation steps and let referees judge whether the edge is real.","headline":"The abstract claims a transformer end-to-end policy beats rules net of costs on 16 CME futures while LSTM does not, but supplies no validation details so the ranking is hard to trust.","tokens_in":2333,"tokens_out":389,"would_cite":false,"duration_ms":26556,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"End-to-end AI policies that map market states to weights outperform simple rules on cross-asset futures portfolios, especially when transaction costs are considered.","keywords":["end-to-end policy","portfolio timing","futures markets","LSTM","transformer","Sharpe ratio","transaction costs","cross-asset allocation"],"falsifier":"Training the LSTM and transformer on a shifted out-of-sample period or with different input features and checking if the ranking over rules still holds.","tokens_in":2622,"feed_emoji":"📈","tokens_out":615,"duration_ms":28808,"temperature":0.7,"pith_summary":"The authors investigate an alternative to the standard forecast-then-optimize approach by training models to directly produce portfolio weights from market states. They apply this to timing across sixteen liquid CME futures contracts using a loss based on the Sharpe ratio. The resulting policies rank higher than equal weighting, risk parity, and momentum strategies on the combined portfolio and in several asset classes. LSTM and transformer models show similar gross performance, but the transformer trades much less and maintains superiority after costs.","feed_headline":"AI timing beats rules in futures, transformers win on costs","feed_subtitle":"On liquid CME contracts, learned policies outperform equal weighting and momentum after transaction costs due to reduced trading.","key_machinery":"End-to-end parametric portfolio policy, implemented via LSTM or transformer networks, that directly outputs weights from market states and is trained by maximizing a differentiable Sharpe ratio.","core_discovery":"Training end-to-end policies on sixteen CME futures with a differentiable Sharpe ratio loss produces models that rank above equal weighting, risk parity, and time-series momentum on the pooled cross-asset portfolio. While LSTM and transformer architectures perform comparably without costs, the transformer creates a stronger policy by generating far lower turnover, allowing it to match or exceed equal weighting through moderate transaction cost levels.","pith_inferences":["These policies might extend to equity or fixed income timing if similar state representations are used.","Testing on post-sample periods after the training data could confirm stability across market regimes.","Alternative loss functions beyond Sharpe ratio may yield policies with different risk profiles.","The reduced turnover suggests scalability to larger position sizes without proportional cost increases."],"forward_implications":["The learned policies demonstrate an advantage in timing-based tilts that contribute to risk and return in diversified portfolios.","Transformer architectures are more suitable for cost-sensitive applications due to significantly lower trading activity compared to LSTMs.","Performance gains are observed in the overall portfolio and multiple sub-asset classes, though not consistently across all.","Direct optimization of portfolio metrics can bypass the need for intermediate return forecasts."],"fun_headline_variants":["End-to-end policies top rules in liquid CME futures","Transformers match equal weighting through moderate costs","Learned policies top rules on pooled portfolio","Transformer trades less than LSTM with transaction costs"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The out-of-sample advantage of the end-to-end policies over rules is robust to variations in training window, feature construction, and hyperparameter selection.","fun_headline_variants_meta":{"raw":{"variants":["End-to-end policies top rules in liquid CME futures","Transformers match equal weighting through moderate costs","Learned policies top rules on pooled portfolio","Transformer trades less than LSTM with transaction costs"]},"model":"grok-4.3","cost_usd":0.010194,"raw_usage":{"total_tokens":4496,"prompt_tokens":622,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":101937000,"prompt_tokens_details":{"text_tokens":622,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3819,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":622,"tokens_out":55,"duration_ms":39657,"temperature":1.0,"reasoning_tokens":3819,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-02T02:09:33.264855+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Training the LSTM and transformer on a shifted out-of-sample period or with different input features and checking if the ranking over rules still holds.","supporting_citations":[],"review_version":1}