{"id":"0475a203-5799-4088-a88a-b7363b1a405a","arxiv_id":"2605.23959","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces a paired one-switch benchmark that quantifies protocol-induced inflation from decision-time leakage in financial ML backtests on equity panels from 2016-2024.","lead":"The paper introduces a benchmark called When Alpha Disappears that measures how much financial backtest performance inflates due to decision-time data leakage by changing one evaluation rule at a time around a clean reference. A smart generalist might read it to understand common pitfalls that make machine-learning trading strategies appear more profitable in tests than they would be in live trading.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flags the core methodological assumption. Because the paper is a diagnostic benchmark with zero parameters and no formal claims beyond measured differences, the toggling protocol is the only plausible load-bearing point; absent evidence of implementation leakage in the description, the finding stands as reported.","tokens_in":1738,"tokens_out":242,"duration_ms":14276,"concrete_test":"Re-run the full toggle suite on the two equity panels using an independent walk-forward implementation that enforces strict temporal separation at the feature-construction step; compare the magnitude and ranking of inflation effects to the original tables.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical observation from a controlled one-at-a-time toggle benchmark around a t+1-open reference. The methodology description contains no equations, derivations, or modeling assumptions that can be checked for hidden inconsistencies; the selectivity result is presented as a measured outcome rather than a derived necessity. The weakest_assumption identified by the reader (isolated toggling while holding panel/split/model/horizon fixed) is the intended design and does not appear violated by the stated protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces the 'When Alpha Disappears' benchmark, which diagnoses decision-time leakage in financial ML backtests via one-at-a-time toggles of evaluation conventions around a fixed clean t+1-open reference (holding data panel, walk-forward split, model family, horizon, portfolio rule, and costs fixed). Across two daily-OHLCV equity panels, six model families, and yearly tests 2016-2024, it reports that inflation is highly selective: centered temporal features and same-day-open execution with post-open daily-bar information produce large, stable gains in predictive and trading metrics, while global normalization, future-informed graph structure, and same-day-close execution are weak in most settings. The benchmark is positioned as diagnostic rather than a source of tradable alpha.","tokens_in":1785,"tokens_out":564,"duration_ms":20408,"significance":"If the selectivity results hold under the controlled protocol, the benchmark supplies a practical, reproducible method for quantifying protocol-induced inflation in quant-finance backtests. The design's emphasis on isolated toggles, multiple panels/models/periods, and explicit reference case is a strength that could help standardize evaluation practices and reduce over-optimism in the field.","major_comments":[{"comment":"Methods (or equivalent section describing the benchmark protocol): the claim that toggles are performed while 'holding the data panel, walk-forward split, model family, horizon, portfolio rule, and cost convention fixed' requires explicit pseudocode or a table enumerating the exact feature-construction and execution rules for the t+1-open reference versus each toggle; without this, readers cannot verify that no unintended leakage was introduced during the 'clean' baseline construction.","section":"Methods"},{"comment":"Results section (tables or figures reporting metric changes): the statements of 'large and stable increases' and 'weak in most settings' need accompanying effect-size tables (e.g., mean and std of Sharpe or accuracy deltas across the 9 years) and a clear statement of the statistical test used to classify an effect as 'large' versus 'weak'; the current description supplies no raw numbers or significance thresholds, which is load-bearing for the selectivity conclusion.","section":"Results"}],"minor_comments":[{"comment":"The abstract states findings but omits any numerical illustration of the reported deltas; adding one or two concrete effect sizes would improve immediate readability.","section":"Abstract"},{"comment":"Notation for the six model families and two equity panels should be defined at first use (or in a table) rather than left implicit.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments, which will improve the clarity and reproducibility of the manuscript. We address each major comment below.","responses":[{"response":"We agree that the current manuscript lacks sufficient detail on the exact rules. In the revised version, we will add a dedicated table (or pseudocode block) in the Methods section that explicitly enumerates the feature-construction steps, execution timing, and information sets for the t+1-open reference case and for each of the toggled variants. This will allow readers to verify the isolation of each toggle.","revision_made":"yes","referee_comment":"[Methods] Methods (or equivalent section describing the benchmark protocol): the claim that toggles are performed while 'holding the data panel, walk-forward split, model family, horizon, portfolio rule, and cost convention fixed' requires explicit pseudocode or a table enumerating the exact feature-construction and execution rules for the t+1-open reference versus each toggle; without this, readers cannot verify that no unintended leakage was introduced during the 'clean' baseline construction."},{"response":"We accept that quantitative support for the selectivity claims is required. The revised manuscript will include new tables reporting, for each toggle, the mean and standard deviation of deltas in key metrics (Sharpe ratio, accuracy, etc.) across the nine yearly periods. We will also state the statistical procedure (paired t-test on yearly deltas, with effect-size thresholds) used to classify effects as large versus weak, including any multiple-testing adjustments.","revision_made":"yes","referee_comment":"[Results] Results section (tables or figures reporting metric changes): the statements of 'large and stable increases' and 'weak in most settings' need accompanying effect-size tables (e.g., mean and std of Sharpe or accuracy deltas across the 9 years) and a clear statement of the statistical test used to classify an effect as 'large' versus 'weak'; the current description supplies no raw numbers or significance thresholds, which is load-bearing for the selectivity conclusion."}],"tokens_in":1420,"tokens_out":444,"duration_ms":14288,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is a benchmark that flips one evaluation rule at a time around a clean t+1-open reference to measure how much each choice inflates backtest numbers. They hold the data panel, splits, models, horizon, and costs fixed while testing across two daily equity sets, six model families, and 2016-2024 yearly windows.\n\nWhat is new is the paired toggle design itself. It turns leakage from a vague worry into something measured by changing single conventions like feature centering, execution timing, normalization, and graph structure. The results show clear selectivity: centered temporal features and same-day-open execution with post-open bar data produce large, stable lifts in both predictive and trading metrics. Global normalization, future-informed graphs, and same-day-close execution stay weak in most runs.\n\nThe paper does well at keeping the experiment diagnostic rather than promotional. It does not claim new alpha and focuses on making protocol fragility visible. The controlled setup matches the stated goal and avoids obvious circularity.\n\nSoft spots are mostly about missing detail. The abstract supplies no implementation code, exact statistical tests, or variance numbers, so the size and stability of the selectivity effects are hard to judge without the full tables. If the differences rest on small panels or untested run-to-run noise, the practical takeaway shrinks. Generalization beyond the tested equities and daily frequency also stays open.\n\nThis is for quant researchers and practitioners who run machine-learning backtests and want to audit their own pipelines. Readers who already care about evaluation standards will find direct, usable checks.\n\nIt deserves a serious referee. The problem is common, the method is straightforward, and the selectivity finding gives concrete guidance even if later work refines the numbers.\n\nI would send it for peer review.","headline":"The paper gives a controlled one-switch benchmark for leakage in quant backtests and reports that inflation hits selectively on a few protocol choices.","tokens_in":2264,"tokens_out":430,"would_cite":false,"duration_ms":21422,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A one-switch benchmark shows that decision-time leakage inflates financial backtest metrics only under specific evaluation conventions.","keywords":["decision-time leakage","financial backtests","evaluation benchmark","protocol-induced inflation","machine learning","equity OHLCV","walk-forward testing"],"falsifier":"Running the same toggles on a third equity panel or with a different walk-forward scheme and finding that the large inflation from centered features and same-day-open execution no longer appears would falsify the selectivity claim.","tokens_in":2620,"feed_emoji":"📉","tokens_out":714,"duration_ms":18033,"temperature":0.7,"pith_summary":"The paper presents a benchmark that measures protocol-induced inflation in financial machine-learning backtests by changing one evaluation rule at a time from a clean t+1-open reference while keeping data, splits, models, and costs fixed. It finds the inflation effect is selective: centered temporal features and same-day-open execution with post-open bar information produce large, stable gains in both predictive accuracy and trading returns, but global normalization, future-informed graphs, and same-day-close execution show little effect across the tested panels and models. A sympathetic reader would care because this isolates which common backtest choices can make models appear profitable when they would not be in live trading without those choices.","feed_headline":"One-switch test finds leakage inflates only certain backtest rules","feed_subtitle":"Toggling conventions around a clean t+1-open reference shows large metric gains from centered features and same-day-open execution, but not","key_machinery":"The one-switch benchmark that isolates decision-time leakage by toggling a single evaluation convention around a clean t+1-open reference while fixing all other protocol elements.","core_discovery":"The benchmark estimates protocol-induced inflation by toggling one evaluation convention at a time around a clean t+1-open reference, while holding the data panel, walk-forward split, model family, horizon, portfolio rule, and cost convention fixed. Across two daily-OHLCV equity panels, six model families, and yearly tests from 2016--2024, inflation is highly selective: centered temporal features and same-day-open execution with post-open daily-bar information cause large and stable increases in both predictive and trading metrics, whereas global normalization, future-informed graph structure, and same-day-close execution are weak in most settings.","pith_inferences":["Backtest results that rely on centered features or same-day-open post-open information may drop sharply once those conventions are removed.","The benchmark could be applied to other asset classes or higher-frequency data to test whether the same selective pattern holds.","Published financial ML papers using the inflating conventions may overstate robustness unless they also report the clean-reference results."],"forward_implications":["Centered temporal features produce large predictive and trading metric gains when toggled on.","Same-day-open execution that includes post-open daily-bar information inflates both predictive and trading metrics in a stable way.","Global normalization produces only weak inflation in most tested settings.","Future-informed graph structures and same-day-close execution show weak effects on metric inflation across the panels and models."],"fun_headline_variants":["One-switch benchmark isolates leakage to centered features and open execution","One-switch test finds leakage selective to temporal centering and same-day opens","Single switch reveals selective leakage in specific backtest rules","Benchmark toggle shows selective leakage from centered temporal features"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Changing one evaluation convention at a time around the clean reference isolates its leakage effect without hidden interactions from the other fixed elements.","fun_headline_variants_meta":{"raw":{"variants":["One-switch benchmark isolates leakage to centered features and open execution","One-switch test finds leakage selective to temporal centering and same-day opens","Single switch reveals selective leakage in specific backtest rules","Benchmark toggle shows selective leakage from centered temporal features"]},"model":"grok-4.3","cost_usd":0.010467,"raw_usage":{"total_tokens":4631,"prompt_tokens":673,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":104674500,"prompt_tokens_details":{"text_tokens":673,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3894,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":673,"tokens_out":64,"duration_ms":33481,"temperature":1.0,"reasoning_tokens":3894,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T22:36:11.941293+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same toggles on a third equity panel or with a different walk-forward scheme and finding that the large inflation from centered features and same-day-open execution no longer appears would falsify the selectivity claim.","supporting_citations":[],"review_version":1}