{"id":"7b6b2515-f40e-43db-97a1-aab5150655ea","arxiv_id":"2607.09230","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"In Binance BTC/ETH futures event windows, pre-event L2 liquidity state dominates post-event regime prediction; order flow helps only as an ETH-dominant, stress-amplified overlay.","lead":"Pre-event L2 liquidity state, not order flow or event labels, is the main predictor of post-event liquidity regimes in crypto futures. The finding gives a concrete baseline that execution, RL, and LLM market models should beat before claiming added value.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged label-construction risk; the staged OOS evidence for the stated claim is internally consistent.","rationale":"The strongest claim is carefully scoped: within macro-event windows, pre-event L2 state is first-order for a supervised discrete post-event liquidity regime; continuous logits fail; shallow nonlinear L2 adds a comparable robust gain; order flow adds only as overlay, ETH-dominant and stress-amplified, not established for BTC at both horizons. Table 1, the per-symbol/regime splits, event-clustered bootstrap, and blocked permutation tests support that ranking on the chosen target. The single most load-bearing vulnerability is exactly the one the reader named—whether the equal-weight tercile construction is the right discrete target—because every subsequent comparison is measured against or conditioned on it. The paper does not claim universality of the label, only that under this explicit construction the staged increments hold; the design-principle framing is already qualified as a baseline others must exceed. No stronger internal inconsistency (leakage, circular scoring, or untested first-order claim) is evident from the manuscript. Therefore the reader's CONDITIONAL verdict (pending artifacts and broader coverage) already absorbs the residual risk; no adjustment is warranted.","tokens_in":13973,"tokens_out":718,"duration_ms":7802,"concrete_test":"Re-label the same panel with two alternative discretizations (e.g., (i) PCA/k-means or quantiles on continuous multi-scale L2 descriptors only, and (ii) a continuous liquidity-withdrawal index as in related work [18]), re-run the full staged sequence of Section 4 (coarse baseline → logits → shallow nonlinear L2 → flow overlay) under identical rolling folds and flow-shuffle nulls; if the coarse-state first-order ranking or the ETH-stress / BTC-not-established pattern reverses or loses significance, the label-construction concern lands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly isolates the softest point: the hand-built three-level state (equal-weight count of oriented relative spread, top-20 depth, and top-20 imbalance in train-fold-only top terciles, capped at two; Section 3.2) is the target that defines both the coarse baseline and the post-event label. If that discretization discards the dimensions of liquidity that trade flow actually moves, the state-first ranking and BTC non-establishment could be label artifacts. However, the paper already supplies partial internal checks that keep this from being a silent failure: the joint state beats any single descriptor (0.048 vs 0.033 at 5m), beats a train-fold realized-vol tercile (excess 0.026/0.029 with intervals above zero), continuous logits fail while shallow nonlinear L2-shape recovers a robust gain of comparable size, and order-flow is tested only as an overlay that must clear a flow-shuffle null blocked by month/symbol/pre-event state (and hour). Those controls make the ranking internally coherent for the stated target; they do not prove the target is the only or best liquidity definition. That residual is a scope/label-sensitivity limit already reflected in the CONDITIONAL verdict, not a new load-bearing inconsistency that overturns the staged claim as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies supervised one-step prediction of a discrete post-event L2 liquidity regime (calm/mixed/stressed) for Binance BTCUSDT and ETHUSDT perpetual futures around scheduled macro announcements (2023–mid-2026). The state is built from relative spread, top-20 depth, and top-20 imbalance via train-fold-only terciles and an equal-weight top-tercile count capped at two. Using rolling monthly OOS folds, event-clustered bootstraps, and blocked permutation tests, the authors stage models so each layer is admitted only if it improves on the layer below on the same panel: a coarse pre-event state baseline strongly beats the marginal baseline; continuous multinomial and ordered logits fail to improve on that state; a shallow nonlinear L2-shape model adds a further robust gain of comparable size; and order-flow features add value only as an overlay on L2 shape, robustly for ETH (largest under pre-event stress) but not established for BTC at both horizons. Macro labels locate windows and matched non-event controls but are not used as predictors. The authors propose a state-first design principle and an evaluation baseline that RL, execution, or LLM context layers should exceed.","tokens_in":14390,"tokens_out":1251,"duration_ms":12987,"significance":"If the staged OOS ranking holds, the paper supplies a concrete, falsifiable baseline and protocol for event-window microstructure prediction that cleanly separates persistent L2 state from order-flow overlays and from event-label content. The evaluation design—train-fold-only thresholds, event-clustered resampling, flow-shuffle nulls (including hour-blocked), per-symbol and per-regime reporting, and dual proper scores—is unusually disciplined for this literature and is itself a reusable contribution. The ETH-dominant, stress-amplified order-flow result is a useful within-asset, regime-conditional finding that prior LOB-ML (mostly price targets), generative state-dependent Hawkes work, and macro-event microstructure studies leave open. Scope is deliberately limited (two assets, one-minute top-20 snapshots, liquidity-state rather than price or P&L target), which keeps the claims proportionate.","major_comments":[{"comment":"Section 3.2 (and the target in Eq. 1): the central ranking is defined relative to a hand-built three-level state (equal-weight count of oriented spread/depth/imbalance in train-fold top terciles, capped at two). The paper already shows the joint state beats any single descriptor and a realized-vol tercile, and that continuous logits fail while shallow nonlinear L2 recovers a robust gain. Those checks make the ranking coherent for this label, but they do not establish label robustness. A load-bearing sensitivity is still missing: re-run the staged sequence under at least one alternative discretization (e.g., different aggregation weights, PCA/quantile of the three descriptors, or a continuous liquidity score binned differently) and report whether the state-first ordering and the ETH-vs-BTC order-flow split survive. Without that, the design principle remains tied to one particular target c","section":null},{"comment":"Sections 5.2–5.3 and Table 1 / Figure 2: the order-flow claim is correctly stated as ETH-dominant and not established for BTC, but the manuscript still leans on a pooled overlay that clears its null only because ETH carries it. The operative rows are already the per-asset ones; the abstract and conclusion should lead with the asset- and regime-conditional statement and treat the pooled number as secondary (or drop it from the headline comparison) so that a reader cannot misread a cross-symbol order-flow result that the paper itself does not claim.","section":null}],"minor_comments":[{"comment":"Table 1 note: the scope difference between held-out event-window point estimates and full-panel cluster-bootstrap intervals is carefully disclosed but easy to miss; a one-sentence reminder in the table caption would help.","section":null},{"comment":"Section 4.1: the shallow GBM hyperparameters (depth 3, 60 iterations, lr 0.05, L2=1.0) are stated once; repeating them in the Table 1 caption or a short methods box would make the “shallow nonlinear” claim fully self-contained.","section":null},{"comment":"Figure 2: axis labels and null thresholds are clear, but adding the numerical joint increments (or null 95th percentiles) on or beside the bars would reduce reliance on the prose for the stress-amplification claim.","section":null},{"comment":"Section 6: the explicit non-claims (no P&L, no event-label causality, no sub-second execution, two-asset limit) are well placed; a single sentence cross-referencing the data-resolution limits of Section 3.1 would further protect against over-reading.","section":null},{"comment":"References: several arXiv preprints are recent and relevant; ensure final versions or DOIs are updated at production if available.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a solid empirical contribution with an unusually careful evaluation protocol for q-fin.TR. The residual risk is label sensitivity of the hand-built state, not internal inconsistency or circularity; requiring one alternative discretization as a revision check is proportionate and should not reopen the whole design. Fit for a microstructure / market-design venue is good; the two-asset crypto-futures scope is a genuine limit but is disclosed. I would not block on novelty relative to latent-regime or price-target LOB work—the supervised transition task and staged admission protocol are distinct enough."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: within scheduled macro windows on Binance BTC and ETH futures, the pre-event discrete L2 liquidity state (calm/mixed/stressed from spread, top-20 depth, imbalance) is the main predictor of the post-event regime. Continuous logits fail to beat a coarse state baseline; a shallow nonlinear L2-shape model adds a further robust gain of similar size; order flow helps only as an overlay, clearly for ETH (largest under stress) and not established for BTC at both horizons.\n\nWhat is new is not liquidity regimes or order-flow informativeness—both are well cited—but the combination: a supervised pre-to-post discrete transition task under event windows, a staged layer-admission protocol (each layer must beat the one below on the same panel), and the within-asset, regime-conditional finding that order-flow additivity for this target is asset- and state-dependent. The evaluation is stronger than typical LOB ML work: rolling monthly OOS, train-fold-only terciles, event-clustered bootstraps, blocked flow-shuffle nulls (including hour), NLL and Brier agreement, and honest per-symbol/regime reporting. Table 1 and Figure 2 line up with the prose. Related-work placement is clean; they do not overclaim against price-prediction or latent-regime papers.\n\nSoft spots are real but proportionate. The hand-built three-level state is the load-bearing label; if it discards dimensions that flow actually moves, rankings could be label artifacts. The paper already checks that the joint state beats single descriptors and a vol tercile, and that continuous logits fail while nonlinear shape recovers, so the ranking is coherent for this target—it is not a silent failure. Two-asset scope, one-minute top-20 resolution, and unreleased code/data keep the design-principle framing a bit broader than the evidence. Those are scope limits, not internal contradictions.\n\nThis is for people building event-conditioned microstructure, execution, or RL baselines who need a concrete bar before crediting extra layers. It deserves a serious referee. I would engage with it and cite the protocol and the ETH/BTC asymmetry when working on similar tasks.","headline":"Careful staged OOS evidence that pre-event L2 liquidity state is first-order for discrete post-event regimes on Binance futures, with order-flow value only as an ETH-dominant overlay.","tokens_in":15021,"tokens_out":562,"would_cite":true,"duration_ms":6291,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Within crypto futures event windows, pre-event L2 liquidity state is the first-order predictor of the post-event regime; order flow adds value only as an overlay, and only robustly for ETH under stress.","keywords":["market microstructure","limit order book","liquidity state","order flow","crypto futures","out-of-sample evaluation","state-dependent prediction"],"falsifier":"Rebuild the same staged comparison after replacing the three-level tercile state with an alternative liquidity target (continuous liquidity-withdrawal index, latent regime detector, or different descriptor set) and check whether continuous L2 or order-flow layers then beat the coarse pre-event state at both horizons for both assets.","tokens_in":14832,"feed_emoji":"📊","tokens_out":789,"duration_ms":7381,"temperature":0.7,"pith_summary":"The paper asks what actually predicts the next liquidity regime of a crypto futures book around scheduled macro releases. Using top-20 order-book snapshots and trade flow for Binance BTCUSDT and ETHUSDT, it builds a supervised three-class task: predict whether the post-event book is calm, mixed, or stressed from information available before the release. A staged out-of-sample protocol admits each feature layer only if it beats the layer below it. The first-order signal is simply the pre-event liquidity state itself; continuous linear models over the same book features do not improve on that coarse baseline, while a shallow nonlinear model of book shape adds a further robust gain of similar size. Order flow then helps only when layered on top of that book model, never as a replacement, and the help is asset- and state-dependent: clear and stress-amplified for ETH, not established for BTC at both horizons. The practical claim is a state-first design rule: any richer execution, reinforcement-learning, or language-model context layer should first beat this liquidity-state transition baseline before its added value is credited.","feed_headline":"Pre-event book state, not order flow, leads crypto liquidity transitions","feed_subtitle":"On Binance futures, order flow helps only as an overlay—and mainly for ETH under stress.","key_machinery":"The supervised discrete L2 liquidity-state transition task: a three-level calm/mixed/stressed label built from train-fold-only terciles of oriented relative spread, top-20 depth, and top-20 imbalance, evaluated in a staged sequence under rolling monthly out-of-sample folds, event-clustered resampling, and blocked permutation tests that admit each feature layer only if it improves on the layer below it.","core_discovery":"Inside scheduled macro-event windows on Binance BTC and ETH perpetual futures, the primary predictor of the discrete post-event L2 liquidity regime is the pre-event L2 liquidity state. A coarse pre-event state baseline strongly improves on a marginal baseline; multinomial and ordered logits over continuous L2 features fail to improve on that state; a shallow nonlinear L2-shape model adds a further robust gain of comparable size; and local order flow contributes incremental predictive value only as an overlay on the L2 model, robustly for ETH (largest under stressed pre-event liquidity) but not established for BTC across both horizons.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Pre-event L2 state drives crypto liquidity regimes more than order flow","Book state baseline beats continuous L2 features in event windows","Order flow helps only as L2 overlay—robust for ETH under stress","State-first: pre-event liquidity predicts post-event crypto regimes","L2 pre-event state leads; flow adds value only on top for ETH"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The hand-built three-level liquidity state from equal-weight top-tercile counts of spread, depth, and imbalance is the right discrete target; if that label discards the dimensions of liquidity that trade flow actually moves, the state-first ranking and the BTC non-result could be artifacts of the discretization.","fun_headline_variants_meta":{"raw":{"variants":["Pre-event L2 state drives crypto liquidity regimes more than order flow","Book state baseline beats continuous L2 features in event windows","Order flow helps only as L2 overlay—robust for ETH under stress","State-first: pre-event liquidity predicts post-event crypto regimes","L2 pre-event state leads; flow adds value only on top for ETH"]},"model":"grok-4.5","effort":"low","cost_usd":0.005768,"raw_usage":{"total_tokens":1628,"prompt_tokens":952,"num_sources_used":0,"completion_tokens":97,"cost_in_usd_ticks":57680000,"prompt_tokens_details":{"text_tokens":952,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":579,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":952,"tokens_out":97,"duration_ms":6138,"temperature":1.0,"reasoning_tokens":579,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T04:31:45.948830+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rebuild the same staged comparison after replacing the three-level tercile state with an alternative liquidity target (continuous liquidity-withdrawal index, latent regime detector, or different descriptor set) and check whether continuous L2 or order-flow layers then beat the coarse pre-event state at both horizons for both assets.","supporting_citations":[],"review_version":1}