{"id":"38268ae8-1b6a-4455-b503-1e0294a42fff","arxiv_id":"2603.23571","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Stateful training of a linear-attention model, preserving recurrent memory across training segments, improves long-horizon memory and in-context adaptation on MAZE and ProcTHOR navigation.","lead":"StateLinFormer trains a linear-attention navigator by carrying memory states across batch boundaries instead of resetting them. If the claim holds, long-horizon embodied agents could keep adapting from experience without fixed context windows or hand-built maps.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified — provided full text is a different paper, so the abstract claim cannot be stress-tested on its own terms.","rationale":"The reader's UNVERDICTED / LOW-confidence stance is the only defensible outcome given the materials: abstract-only for the claimed paper, full text of a different paper. No load-bearing technical objection can be formulated against StateLinFormer's central claim because the supporting derivations, ablations, and numbers are absent. Manufacturing a critique from the wrong manuscript would violate good-faith reading. The concrete test simply restates the minimal verification that would allow a real second-pass review once the correct PDF is supplied. Verdict and agreement therefore remain unchanged.","tokens_in":5183,"tokens_out":410,"duration_ms":4939,"concrete_test":"Obtain the true StateLinFormer PDF (2603.23571). Locate the stateful-training procedure and any ablation that freezes or reinitializes memory at segment boundaries while matching compute and context length; recompute the long-horizon MAZE/ProcTHOR curves. If the stateful advantage disappears under matched optimization, the infinite-sequence claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review target is StateLinFormer (arXiv 2603.23571). The only usable material for that paper is its abstract. The CACHEABLE full manuscript is Dual-Criterion Curriculum Learning (arXiv 2603.23573), an unrelated work on hybrid loss/density curricula for time-series forecasting. There is therefore no method section, equation, ablation, or table against which to probe whether carrying recurrent states across training-segment boundaries is a faithful approximation of infinite-sequence learning, or whether reported gains on MAZE/ProcTHOR reflect genuine long-horizon retention versus altered optimization dynamics. The reader's weakest_assumption correctly flags the uncheckable transfer claim; without the actual paper body that assumption cannot be elevated into a concrete technical flaw or dismissed.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The submission under review is titled StateLinFormer and, from its abstract, claims a linear-attention navigation architecture trained with a stateful memory mechanism that preserves recurrent states across consecutive training segments rather than reinitializing them at batch boundaries. This is presented as an approximation to learning on infinitely long sequences, yielding long-horizon memory retention and improved context-dependent adaptation (framed as enhanced in-context learning) on MAZE and ProcTHOR relative to a stateless linear-attention counterpart and fixed-context Transformer baselines. However, the full manuscript text supplied for review is an entirely different paper—Dual-Criterion Curriculum Learning (DCCL) for temporal forecasting (arXiv:2603.23573)—with no method section, equations, architecture description, training algorithm, ablations, or experimental tables for StateLinFormer. The central empirical and methodological claims of StateLinFormer therefore cannot be assessed from the provided materials.","tokens_in":5385,"tokens_out":903,"duration_ms":18881,"significance":"If the abstract claims held under proper evaluation—stateful cross-segment training of linear attention producing genuine long-horizon retention and growing gains with interaction length on standard navigation suites—the work would be of clear interest to the navigation, sequence modeling, and in-context learning communities, offering a practical alternative to modular mapping systems and fixed-window Transformers. That significance remains conditional: without the actual StateLinFormer manuscript body, no credit can be assigned for soundness, baselines, ablations, or reproducibility, and the contribution cannot be ranked against prior stateful or recurrent-memory training literature.","major_comments":[{"comment":"Manuscript identity mismatch: the title, abstract, and arXiv id (2603.23571, StateLinFormer) do not match the full text provided, which is Dual-Criterion Curriculum Learning for temporal data (arXiv:2603.23573). There is no StateLinFormer architecture, stateful training algorithm, loss, or navigation experiment in the body. The central claims (approximation to infinite-sequence learning; outperformance on MAZE/ProcTHOR; ICL gains with interaction length) are therefore unverifiable. This is load-bearing: a referee cannot evaluate soundness, baselines, or the weakest transfer assumption without the correct paper.","section":null},{"comment":"Abstract-only claim of approximating infinite-sequence learning by preserving recurrent states across training-segment boundaries cannot be checked. No derivation, pseudocode, or ablation is available to distinguish genuine long-horizon retention at test time from altered optimization dynamics or short-horizon fitting. Until the correct method section and ablations (stateful vs. stateless under matched compute/context, state carry-over vs. longer segments, transfer to held-out horizons) are supplied, the main scientific claim remains unsupported.","section":null},{"comment":"Reported significant outperformance on MAZE and ProcTHOR, and the claim that gains grow with interaction length as evidence of enhanced ICL, cannot be inspected: no tables, metrics, error bars, baseline definitions, or context-window controls appear in the supplied text. These results are load-bearing for the paper’s contribution and must be present and reproducible before acceptance can be considered.","section":null}],"minor_comments":[{"comment":"Once the correct StateLinFormer manuscript is provided, ensure the abstract’s ICL language is tied to a precise operational definition (e.g., adaptation curves vs. interaction length under fixed parameters) rather than left as a suggestion.","section":null},{"comment":"Editorial/process note: the cacheable full-text prefix for this review pass contains the wrong arXiv paper; the authors or submission system should re-upload 2603.23571 so that section-, equation-, and table-level review is possible.","section":null}],"recommendation":"uncertain","confidential_remarks":"The supplied full manuscript is unambiguously a different paper (DCCL / 2603.23573). This looks like a pipeline or packaging error rather than author misconduct, but it makes a normal technical review of StateLinFormer impossible. I recommend returning the submission for the correct PDF/source and re-review; I would not treat the current package as a complete submission. Confidence in any scientific judgment on StateLinFormer is necessarily low until the real body is available."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The only usable material for 2603.23571 is the abstract. The full text we were given is Dual-Criterion Curriculum Learning (2603.23573), an unrelated time-series curriculum paper. So we cannot check methods, ablations, tables, or the central transfer claim.\n\nWhat the abstract actually asserts is an applied training recipe: linear-attention navigation (StateLinFormer) that keeps recurrent memory states across consecutive training segments instead of resetting them at batch boundaries, framed as approximating infinite-sequence learning. Claimed wins are over a stateless linear-attention twin and fixed-window Transformers on MAZE and ProcTHOR, with gains that grow with interaction length and are read as better in-context adaptation. That is a legitimate embodied-AI angle if the numbers hold. Linear attention and carrying state across segments are not new ideas; the contribution, if any, is the navigation packaging and the length-dependent ICL story.\n\nSoft spots are not subtle technical flaws in a body we do not have; they are absence of evidence. We cannot see whether state carry-over is a faithful long-horizon approximation or just changes optimization dynamics, whether baselines and metrics are fair, or whether error bars and ablations support the ICL interpretation. The reader’s weakest assumption is exactly right and currently uncheckable. Citation pattern and math cannot be assessed from an abstract alone.\n\nWho this is for: people working on long-context policies and end-to-end navigation who already track linear-attention and stateful training tricks. Until the real manuscript is attached, it does not deserve reading-group time or a citation commitment from me. A serious editor should not send this package to referees as-is; get the correct PDF first. If the real paper later shows clean length-scaling ablations and honest baselines, re-open it. Right now: do not engage.","headline":"Wrong paper body was supplied for StateLinFormer; only the abstract is usable, so the long-horizon/ICL claims cannot be audited and this is not ready for a serious read.","tokens_in":5980,"tokens_out":468,"would_cite":false,"duration_ms":5158,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Training a linear-attention navigator by carrying memory states across batch boundaries yields long-horizon retention that fixed-context Transformers lack.","keywords":["stateful training","linear attention","long-term memory","navigation","in-context learning","recurrent memory states","fixed context windows"],"falsifier":"Train the identical linear-attention architecture with and without state carry-over, then measure navigation success and context-dependent adaptation on held-out episodes whose length far exceeds any single training segment; if the stateful version shows no growing advantage (or worse performance) as horizon increases, the central claim fails.","tokens_in":6079,"feed_emoji":"🧭","tokens_out":765,"duration_ms":11576,"temperature":0.7,"pith_summary":"Navigation agents need lasting memory: both to generalize immediately from context and to keep adapting over long interactions. Modular map-based systems are rigid, while standard Transformers forget once the fixed context window rolls over. This paper claims that a linear-attention model, trained so its recurrent memory states are kept alive across consecutive training segments instead of being wiped at every batch boundary, effectively learns as if sequences were infinitely long. On maze and household navigation benchmarks the resulting StateLinFormer beats both its own stateless linear-attention twin and ordinary fixed-window Transformers, and the gap widens as episodes get longer, which the authors read as stronger in-context learning for navigation.","feed_headline":"Stateful training gives navigators long-horizon memory","feed_subtitle":"Keeping memory alive across training batches beats fixed-window Transformers as episodes lengthen","key_machinery":"Stateful training for linear attention: the recurrent memory state produced by one training segment is passed unchanged as the initial state of the next segment instead of being zeroed at the batch boundary, so gradient updates see a continuous memory trajectory.","core_discovery":"Preserving recurrent memory states across training-segment boundaries, rather than reinitializing them, lets a linear-attention navigation model approximate learning on unbounded sequences and thereby retain usable long-horizon memory at test time, outperforming both stateless linear attention and fixed-context Transformers, with larger gains as interaction length grows.","pith_inferences":["The same cross-segment state carry-over could be tried on other recurrent or linear-attention sequence models outside navigation (e.g., long dialogue or continual control).","If the benefit is mainly better optimization rather than true infinite-context learning, shorter but carefully scheduled segments might recover most of the gain at lower cost.","Explicit comparison against map-based modular agents under identical long-horizon protocols would clarify whether end-to-end stateful memory closes the flexibility gap the abstract highlights."],"forward_implications":["Long-horizon navigation can be improved without enlarging the attention window or building an explicit map module.","Linear-attention architectures become competitive for persistent memory once training itself is made stateful.","In-context adaptation for sequential decision tasks strengthens when training no longer resets internal state at every batch.","As episode length grows, the relative benefit of stateful training over fixed-context baselines is expected to increase."],"fun_headline_variants":["Stateful training preserves memory across batches for navigation","Linear navigators gain long-horizon recall via unreset states","Persistent training states beat fixed-window Transformers on long episodes","StateLinFormer retains usable memory by never reinitializing batches","Keeping states alive across segments lifts navigation ICL with length"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Carrying memory states across artificial training cuts is assumed to be a faithful stand-in for true infinite-horizon learning, and the resulting states are assumed to transfer into genuine long-term retention rather than merely changing short-horizon optimization dynamics.","fun_headline_variants_meta":{"raw":{"variants":["Stateful training preserves memory across batches for navigation","Linear navigators gain long-horizon recall via unreset states","Persistent training states beat fixed-window Transformers on long episodes","StateLinFormer retains usable memory by never reinitializing batches","Keeping states alive across segments lifts navigation ICL with length"]},"model":"grok-4.5","effort":"low","cost_usd":0.007004,"raw_usage":{"total_tokens":1662,"prompt_tokens":694,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":70040000,"prompt_tokens_details":{"text_tokens":694,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":885,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":694,"tokens_out":83,"duration_ms":8400,"temperature":1.0,"reasoning_tokens":885,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T19:53:32.816188+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train the identical linear-attention architecture with and without state carry-over, then measure navigation success and context-dependent adaptation on held-out episodes whose length far exceeds any single training segment; if the stateful version shows no growing advantage (or worse performance) as horizon increases, the central claim fails.","supporting_citations":[],"review_version":1}