{"id":"c1511154-fd74-4730-b2d2-c31074ae6fa8","arxiv_id":"2405.11597","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"PredFT improves fMRI-to-text decoding by using a side network with self-attention to extract brain predictive representations based on predictive coding theory and fusing them into the main network, outperforming baselines on two datasets.","lead":"The paper proposes PredFT, a neural model that adds a side network using self-attention on brain ROIs to capture predictive representations from fMRI signals and fuse them into language decoding. A smart generalist might read it for progress toward more accurate brain-to-text systems that could aid communication for those who cannot speak.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Self-attention on ROIs lacks explicit mechanism to enforce multi-timescale future prediction, so outperformance may not stem from predictive coding","rationale":"The reader's weakest assumption identifies the same methodological gap that bears on the outperformance claim. Full text availability does not alter this because the abstract already isolates the unverified step; any additional implementation details would need to be checked against the concrete ablation above before the empirical result can be attributed to predictive coding.","tokens_in":1663,"tokens_out":300,"duration_ms":16104,"concrete_test":"Ablate the side network entirely (or replace self-attention with a simple linear projection of the same ROIs) and retrain on both datasets; if BLEU/ROUGE or other reported metrics drop by less than the original margin over baselines, the predictive-representation claim does not explain the gains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (PredFT outperforms baselines on two fMRI datasets) rests on the side network producing brain predictive representations that meaningfully improve continuous decoding. The abstract states only that a self-attention module is applied to related ROIs and the result is fused; no loss term, future-word prediction objective, or timescale-specific architecture is described. Without these, the side network reduces to an auxiliary attention block whose benefit could arise from extra capacity rather than predictive-coding alignment. This directly weakens the link between the reported gains and the motivating theory.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes PredFT, a two-network architecture for fMRI-to-text decoding that incorporates predictive coding theory: a main network performs continuous language reconstruction while a side network applies self-attention over related ROIs to produce 'brain predictive representations' that are fused into the main network. Experiments on two naturalistic language-comprehension fMRI datasets are reported to show that PredFT outperforms existing decoding models on several (unspecified) evaluation metrics.","tokens_in":1776,"tokens_out":406,"duration_ms":14864,"significance":"If the reported gains are shown to arise specifically from alignment with multi-timescale predictive representations rather than from added capacity alone, the work would supply a concrete architectural bridge between predictive-coding accounts of language comprehension and practical brain-signal decoding, potentially improving robustness of continuous reconstruction.","major_comments":[{"comment":"Model description (side-network paragraph): the side network is said to 'obtain brain predictive representation' via self-attention on ROIs, yet no future-word prediction loss, next-token objective, or explicit multi-timescale regularizer is defined; without such a term the side network reduces to an auxiliary attention block whose benefit cannot be attributed to predictive coding.","section":"Model description"},{"comment":"Experiments section: the central claim that PredFT 'outperforms current decoding models on several evaluation metrics' supplies no baselines, no metric definitions, no statistical tests, no data-split protocol, and no control for parameter count; these omissions make the empirical support for the claim impossible to evaluate.","section":"Experiments"}],"minor_comments":[{"comment":"The abstract states that the side network is 'fused into the main network' but does not specify the fusion operation (concatenation, cross-attention, etc.); this should be stated explicitly with an equation.","section":"Abstract / Model"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the thoughtful review and the opportunity to clarify the manuscript. We address the two major comments point-by-point below and will incorporate revisions where appropriate to strengthen the paper.","responses":[{"response":"We agree that the absence of an explicit prediction loss or multi-timescale regularizer limits the direct attribution of the side network's benefit to predictive coding theory. The self-attention module is intended to capture integrative representations across ROIs hypothesized to encode predictions at varying timescales, but this remains an implicit alignment rather than an explicit objective. We will revise the model description section to more precisely articulate this distinction, acknowledge the limitation, and note that future work could incorporate a next-token prediction loss to strengthen the link. This constitutes a partial revision focused on clarification rather than architectural change.","revision_made":"partial","referee_comment":"[Model description] Model description (side-network paragraph): the side network is said to 'obtain brain predictive representation' via self-attention on ROIs, yet no future-word prediction loss, next-token objective, or explicit multi-timescale regularizer is defined; without such a term the side network reduces to an auxiliary attention block whose benefit cannot be attributed to predictive coding."},{"response":"The referee is correct that the current experimental reporting is insufficient for evaluation. The revised manuscript will expand the Experiments section to explicitly list all baselines, define each evaluation metric, report statistical tests with p-values, detail the data-split protocol (including subject-wise or session-wise partitioning), and include parameter-count-matched controls. These additions will be presented in a new table or subsection for transparency.","revision_made":"yes","referee_comment":"[Experiments] Experiments section: the central claim that PredFT 'outperforms current decoding models on several evaluation metrics' supplies no baselines, no metric definitions, no statistical tests, no data-split protocol, and no control for parameter count; these omissions make the empirical support for the claim impossible to evaluate."}],"tokens_in":1304,"tokens_out":429,"duration_ms":22119,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core move is to split the decoder into a main network and a side network. The side network runs self-attention over related fMRI ROIs to produce a predictive representation that gets fused back in. Experiments on two naturalistic language fMRI datasets show better scores than prior decoding models on several metrics. That is the claim in one sentence.","headline":"PredFT adds a side network with self-attention on ROIs to capture predictive coding for fMRI language decoding and reports gains on two datasets, but the architecture lacks any explicit future-prediction mechanism so the theory link is thin.","tokens_in":2262,"tokens_out":162,"would_cite":false,"duration_ms":13711,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"fMRI predictive-coding decoder with ROI self-attention side-network; no structural overlap with RS forcing chain","alignment":"orthogonal","rationale":"Paper's core machinery (main Transformer decoder + side-network ROI self-attention fused via PC-Attn, joint CE loss on original/predicted words, FIR latency compensation) operates entirely in the domain of brain-signal decoding and predictive-coding neuroscience. RS framework derives spacetime, c=1, ℏ, G, 3D, 8-tick periodicity and J-cost from a single distinction (reality_from_one_distinction, AbsoluteFloorClosure, Cost.FunctionalEquation.washburn_uniqueness_aczel, AlexanderDuality.alexander_duality_circle_linking). No shared primitives, cost functions, ratio symmetry, ladder spacings or parameter-free constant derivations appear; domain is outside RS scope.","tokens_in":62192,"confidence":"high","tokens_out":194,"duration_ms":6962,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A side network using self-attention on brain ROIs to extract predictive representations improves fMRI-to-text decoding.","keywords":["fMRI","language reconstruction","predictive coding","brain decoding","self-attention","text generation","naturalistic datasets","PredFT"],"falsifier":"An ablation study in which the side network is removed or replaced with non-predictive random features and the performance advantage on the two datasets disappears or reverses.","tokens_in":2561,"feed_emoji":"🧠","tokens_out":496,"duration_ms":15063,"temperature":0.7,"pith_summary":"The paper tries to establish that language reconstruction from fMRI signals benefits when the decoder explicitly incorporates the brain's natural tendency to predict upcoming words across multiple timescales. It does this by adding a side network that applies self-attention to related regions of interest to derive predictive brain representations and then fuses those representations into the main decoding network. Experiments on two naturalistic language comprehension datasets show the resulting PredFT model outperforms prior decoding approaches on standard metrics. A sympathetic reader would care because the method supplies a neurological grounding for why certain brain signals help generate fluent text rather than treating signals as static features. If correct, the approach suggests that future decoding systems should treat brain activity as an active prediction process instead of a passive readout.","feed_headline":"Predictive side network lifts fMRI language decoding","feed_subtitle":"Fusing self-attention on brain ROIs improves reconstruction from two naturalistic comprehension datasets.","key_machinery":"The side network that applies a self-attention module to related regions of interest (ROIs) to extract multi-timescale predictive representations from fMRI signals before fusion into the main decoder.","core_discovery":"PredFT consists of a main network for continuous fMRI-to-text decoding and a side network that obtains brain predictive representations from related ROIs via a self-attention module; these representations are fused into the main network. The design follows from predictive coding theory, which holds that the brain continuously predicts future words spanning multiple timescales. On two naturalistic language comprehension fMRI datasets the fused model outperforms current decoding models across several evaluation metrics.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Side network fuses predictive ROI representations for fMRI decoding","Brain predictive coding integrated in fMRI language reconstruction","Self-attention on ROIs provides predictive input to text decoder","PredFT combines main decoding with side predictive representations"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The self-attention module applied to related ROIs successfully extracts multi-timescale predictive representations from fMRI signals in a manner that meaningfully improves continuous language decoding when fused into the main network.","fun_headline_variants_meta":{"raw":{"variants":["Side network fuses predictive ROI representations for fMRI decoding","Brain predictive coding integrated in fMRI language reconstruction","Self-attention on ROIs provides predictive input to text decoder","PredFT combines main decoding with side predictive representations"]},"model":"grok-4.3","cost_usd":0.00786,"raw_usage":{"total_tokens":3567,"prompt_tokens":631,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":78599500,"prompt_tokens_details":{"text_tokens":631,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2875,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":631,"tokens_out":61,"duration_ms":18473,"temperature":1.0,"reasoning_tokens":2875,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-24T00:46:15.783948+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An ablation study in which the side network is removed or replaced with non-predictive random features and the performance advantage on the two datasets disappears or reverses.","supporting_citations":[],"review_version":1}