{"id":"1257a1e1-fdaa-4d8f-aeb3-3d90c2297b65","arxiv_id":"2412.05319","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a probabilistic serial reaction time task, motor imagery performance improved across blocks and was modulated by the last variable event differently than motor execution, suggesting distinct influences.","lead":"This study compared how people improve on a finger-tapping task when they actually tap keys versus when they only imagine tapping. The authors report that the imagining group got faster over time and that different statistical factors drove the two groups' response times.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central Group×Event interaction may be an artifact of unmatched response structure: MI RT adds a left-hand stop response absent in ME, so 'distinct factors' is not yet supported.","rationale":"The paper's strongest claim is that different factors determine Motor Imagery duration and Motor Execution performance. This claim rests entirely on between-group comparisons of reaction times. The most direct threat is that the two groups produce responses with fundamentally different motor structures: ME uses three distinct right-hand keypresses (one per stimulus), MI uses one left-hand spacebar after imagining the movement. Consequently, any interaction between Group and Event may be trivially produced by response selection: in ME, events differ in the finger-specific response required (index for '1', middle for '2', ring for '3'), whereas in MI the spacebar response is identical for all events. The observed ME event effects (e.g., F1 vs F2) are naturally explained by response-selection differences; the absence of such a difference in MI is expected. The reader identified this as the weakest assumption, and I agree. The paper does provide a code link and transparent analyses, but no control equates the response demands. An additional concern is the group difference in KVIQ scores (Table 1: Visual 35.1 vs 26.1, Kinesthetic 32.9 vs 25.8), which suggests MI participants may have poorer imagery ability, but the response-structure confound is more fundamental because it directly undermines the measurement of the dependent variable. The proposed control experiment — having ME also press the left-hand spacebar after the right-hand tap — would make the response structure more comparable and allow the interaction to be interpreted. Until such a control is run or an analytical correction is provided, the central claim should be regarded as conditional, not established. Thus the reader's CONDITIONAL verdict is appropriate, and my recommendation is UNCHANGED.","tokens_in":10833,"tokens_out":6955,"duration_ms":73512,"concrete_test":"Run a control experiment (n per group) in which Motor Execution participants perform the same right-hand finger taps and then press the left-hand spacebar to end each trial, with RT measured as stimulus offset to the spacebar press, exactly parallel to the MI group's response. Re-run the two-way mixed ANOVA on these matched RTs. If the Group×Event interaction in Results §1 disappears or changes sign, the current claim that different factors determine MI vs ME performance is unsupported; if it persists, the response-structure confound is ruled out. Additionally, for the MI group, run a baseline simple-RT task where the same left-hand spacebar is pressed to each auditory stimulus without imagery; subtracting this baseline from MI RTs should preserve the Event and LastVariableEvent effects if the stop response is not the source.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the abstract's claim that imagery duration is influenced by factors distinct from execution, the experiment must isolate imagery duration in the RT measure. In Methods 3b, ME participants respond by pressing one of three right-hand keys (1/2/3), while MI participants imagine the corresponding finger movement and then press a single left-hand spacebar to signal completion. Thus the ME RT comprises stimulus identification plus selection among three finger-specific responses; the MI RT comprises stimulus identification plus imagined movement time plus a single fixed stop response. The significant Group×Event interaction (Results §1, mixed ANOVA) and the different Event×LastVariableEvent patterns (Results §2) could therefore reflect the absence of response-selection demands in MI rather than a genuine difference in the temporal organization of imagery and execution. The Discussion's assertion that RT is an 'indirect measure' does not resolve this: without a condition that equates response structure, the interaction is confounded. This is a specific, testable threat to the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper compares motor execution (ME) and motor imagery (MI) in a serial reaction time task with a probabilistic context-tree sequence of three auditory stimuli. Participants in the ME group (n=10) respond by pressing one of three right-hand keys, while participants in the MI group (n=10) imagine performing the corresponding finger movements and then signal completion with a single left-hand spacebar press. The authors report that MI reaction times decrease across blocks, that both groups are sensitive to event probabilities, and that a significant Group×Event interaction indicates that imagery duration is influenced by factors distinct from those influencing execution RT.","tokens_in":10960,"tokens_out":5617,"duration_ms":55093,"significance":"The study uses an elegant probabilistic sequence design and makes its code available, which is a strength. The statistical approach (rank-transformed mixed ANOVAs with Greenhouse-Geisser corrections) is appropriate for the within-subject comparisons. If the central claim were supported, it would have implications for motor imagery theory and motor emulation theory. However, the current design confounds group membership with response structure, and the small sample and multiple testing limit the strength of the conclusions.","major_comments":[{"comment":"The central claim that 'the duration of the motor imagery ... are influenced by distinct factors than those of Motor Execution' rests on the significant Group×Event interaction (F3,54=7.172, p<0.01) and on the event-specific differences between groups shown in Figure 4. However, the two groups differ not only in the mental operation but also in response mode: ME participants select among three right-hand keys (1/2/3), whereas MI participants make a single fixed left-hand spacebar press after the imagined sequence. The MI RT therefore includes an extra stop-response component and lacks the three-alternative response-selection component present in ME RT. The observed interaction and group differences could thus reflect the absence of response-selection demands in MI rather than a genuine difference in the temporal organization of imagery and execution. The Discussion's characterization of RT as an 'indirect measure' does not rule out this alternative. To support the abstract's claim, the design needs a control condition that equates response structure—for example, an execution group that performs the three finger movements and then emits a separate stop response, or an imagery group that indicates the imagined finger with a keypress.","section":"Methods 3b; Results §1 (Figs 3-4); Discussion"},{"comment":"The manuscript reports a large number of ANOVA tests (per-group event/block analyses, a mixed group×event analysis, and per-group event×last-variable-event analyses) without any correction for multiple comparisons across these tests. With n=10 per group and rank-transformed data, the power of these tests is limited, and the probability of false positives is inflated. No effect sizes are reported, making it difficult to gauge the magnitude of the significant effects. The authors should report effect sizes (e.g., partial eta-squared) for each ANOVA and either correct for the number of tests or explicitly frame the results as exploratory. The significant interaction that supports the central claim would be considerably more convincing if accompanied by an effect size and a sensitivity analysis.","section":"Methods §4; Results §1-2"}],"minor_comments":[{"comment":"The text has a typo: 'spacebarkey' should be 'spacebar key.'","section":"Methods 3b"},{"comment":"The caption says 'depicted in Figure4' when it should refer to Figure 5.","section":"Figure 5 caption"},{"comment":"The familiarization phase is described only for the execution-style response; it is unclear whether MI participants practiced the left-hand spacebar response or imagined the finger movements during this phase. Please clarify the familiarization procedure for the MI group.","section":"Methods 3a"},{"comment":"The statement 'The mean reaction times for V2 and V3 decreased across the blocks' is not tied to specific statistical results; please indicate which post-hoc comparisons support this claim.","section":"Results §1"},{"comment":"The discussion of the lack of Block effect in the ME group attributes this to task simplicity, but no independent measure of task complexity is provided. This claim should be softened or supported.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The confound between group and response mode is the paper's core problem. While this cannot be fixed by reanalysis alone, the authors could realistically add a control condition in a follow-up experiment or substantially reframe the manuscript's claims as within-group findings. Given the journal's scope, I think major revision is appropriate rather than outright rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe genuinely new piece here is the transfer of the context-tree serial reaction time task from motor execution to motor imagery, plus the analysis of how the last variable event modulates RT. That combination is worth having, and the authors deserve credit for shipping their code, data, and supplementary tables, which makes the analyses checkable. For a small n=10/group behavioral study, the statistics are careful: rank-transformed mixed ANOVAs, Greenhouse-Geisser corrections, and Bonferroni post-hocs where they matter. There is no circularity; the claims are inferred from behavior.\n\nThe soft spot is structural. The two groups don't produce comparable RTs. In Motor Execution, RT is the time to a specific right-hand finger response among three alternatives. In Motor Imagery, RT is the time to imagine that finger movement and then press a single left-hand spacebar. So the ME RT includes stimulus identification plus response selection among three keys; the MI RT adds a fixed stop response that carries no event-specific information. The Group×Event interaction and the different Event×LastVariableEvent patterns could therefore reflect the absence of response-selection demands in MI rather than a genuine difference in how imagery and execution track sequence structure. The authors call RT an indirect measure of imagery duration, but that doesn't dissolve the confound; they need a condition that equates response structure (e.g., a no-imagery control with the same left-hand response, or a stop-signal design). Without that, the abstract's claim that imagery duration is influenced by 'distinct factors' is not yet supported.\n\nSecondary issues: the MI group scored lower on the KVIQ, which wasn't statistically controlled; no effect sizes are reported; and many ANOVAs run without overall multiple-testing control. The absence of a block effect in ME is odd, though their ceiling-effect explanation is plausible.\n\nWho is this for: researchers working on motor imagery or implicit sequence learning. It's a solid pilot that introduces a new paradigm and provides open data. I'd send it to peer review, but with the clear expectation that the authors either add the right control or soften the conclusion to what the current design can support.","headline":"The context-tree SRTT in motor imagery is a useful new combination with open code and data, but the central claim that imagery duration is governed by distinct factors is undercut by an unmatched response structure between groups.","tokens_in":11468,"tokens_out":2920,"would_cite":true,"duration_ms":30258,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Motor imagery and motor execution in a probabilistic finger-tapping task are timed by distinct factors, with imagery reaction times improving across blocks and tracking the last unpredictable stimulus while execution reaction times do not.","keywords":["motor imagery","motor execution","serial reaction time task","context tree","probabilistic sequence learning","reaction time","sequence learning","mental practice"],"falsifier":"Run the Motor Imagery group with the variable-event probabilities reversed (V2 at 74% and V3 at 26%) and include a control condition in which the Motor Execution group performs the real tap and then presses the left-hand key to end the trial; the predicted inversion of the V2-after-V2 and V3-after-V3 reaction-time pattern in imagery, or a shift in the execution pattern toward the imagery pattern in the control, would directly test whether the distinct-factors claim is real or an artifact of the response structure.","tokens_in":10606,"feed_emoji":"🧠","tokens_out":10618,"duration_ms":93628,"temperature":0.7,"pith_summary":"This paper studies whether mentally rehearsing a finger-tapping sequence is timed by the same factors as actually performing it. In a serial reaction time task, two groups of ten participants heard 750 auditory cues generated by a probabilistic context tree with fixed and variable transitions; one group tapped the matching key with the right hand, the other imagined each tap and pressed a key with the left hand when the imagined movement finished. The imagery group's reaction times decreased across five blocks, whereas the execution group showed no block effect, and the two groups reacted differently to deterministic versus variable events and to the most recent variable event. The paper's central claim is that the duration of motor imagery, measured indirectly through reaction times, is influenced by distinct factors from those governing motor execution. If true, mental-practice and rehabilitation protocols should not assume imagined movements are paced by the same cues as real ones.","feed_headline":"Imagined finger taps follow different timing cues than real ones","feed_subtitle":"Imagined reaction times improve and track the last unpredictable event; executed times do not.","key_machinery":"The key machinery is the probabilistic context tree that generates the 750-item auditory sequence. The tree defines two fixed events, F1 (1 follows 3) and F2 (2 follows 1), each with 100% probability, and two variable events, V2 (2 follows 2, 26%) and V3 (3 follows 2, 74%), making some transitions predictable and others not. Reaction time to each event is the behavioral readout of how well the participant is tracking the tree, and in the Motor Imagery group the left-hand spacebar press converts the end of the imagined right-hand finger movement into a measurable reaction time. Comparing the four event types across blocks, and conditioning on the last variable event, is what exposes the different timing factors in the two groups.","core_discovery":"The central discovery is that in a task with an intrinsically variable stimulus sequence, motor imagery performance improves across blocks while motor execution performance does not show a block effect, and the two conditions produce different reaction-time profiles across deterministic (F1, F2) and probabilistic (V2, V3) events. In the imagery group, reaction times were shorter when a variable event repeated its own identity (V2 after V2, V3 after V3), a pattern not seen in execution; and the execution group's difference between F1 and F2 disappeared under imagery. The authors treat reaction time as an indirect measure of the duration of the imagined movement and interpret these contrasts as evidence that motor planning during imagery continuously incorporates the last variable event, so the factors timing imagery are different from those timing executed responses.","pith_inferences":["A direct test of the paper's weakest assumption would be to give the Motor Execution group the same single-response structure, having participants actually perform the indicated right-hand tap and then press a left-hand button to end the trial; if the event-dependent pattern changes to resemble the imagery group's, the distinct-factors claim would instead reflect the response structure.","The 26%/74% probabilities attached to V2 and V3 are natural dials for a follow-up: if imagery reaction times follow the probabilities as they are varied, that would confirm that imagery tracks event likelihood; if not, the effect may be driven by event identity rather than probability.","The paper implies that in rehabilitation, imagined movements should be rehearsed in variable, context-dependent sequences rather than fixed rhythms, since imagery appears sensitive to probabilistic context but not to the same timing cues as execution.","One could extend the design to patient populations where sequence learning and motor imagery are both affected, using the context-tree task to separate predictive-timing deficits from execution deficits."],"forward_implications":["If imagery duration is governed by distinct factors, then mental-practice training must be designed around the probabilistic structure of events rather than merely repeating the movement at a fixed pace.","Imagery reaction times improved across blocks in this task, so mental rehearsal alone supports sequence learning even when the sequence contains unpredictable transitions.","The execution group's F1-versus-F2 contrast, absent in imagery, suggests anticipation of an upcoming variable event is tied to effector-specific motor preparation that does not operate when the movement is only imagined.","The imagery group's sensitivity to the last variable event shows predictive context is updated during imagery, allowing imagery-based protocols to probe predictive sequence learning without overt movement."],"supporting_citations":[{"why":"Established the response-time effect of mispredictions in the same context-tree stochastic game, providing the behavioral precedent for using reaction times as a readout of sequence prediction.","marker":"Cabral-Passos et al. (2024)"},{"why":"Introduced the paradigm for studying how humans learn regularities in sequences with unpredictable steps, which the present task adapts.","marker":"Duarte et al. (2019)"},{"why":"Supplied the specific context tree whose probabilities define the fixed and variable events used in the stimulus sequence.","marker":"Uscapi (2020)"},{"why":"Provided the evidence that the duration of imagined and executed movements is correlated, the background against which the paper claims their determining factors differ.","marker":"Papaxanthis et al. (2002)"},{"why":"Provides the motor emulation theory framework that explains how predictive forward models operate during motor imagery.","marker":"Hurst and Boe (2022)"},{"why":"Supports the claim that motor imagery involves predicting the sensory consequences of the imagined movement, the mechanism behind the imagery timing effects.","marker":"Kilteni et al. (2018)"}],"fun_headline_variants":["Imagined tapping uses different timing rules than real tapping","Motor imagery and execution: distinct timing factors revealed","In imagined tasks, reaction times improve and track unpredictable cues","Imagined finger movements respond to different cues than executed ones"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that the time from the auditory cue to the left-hand spacebar press in the imagery group truly measures how long the imagined right-hand finger movement took, rather than reflecting the different response structure or the time needed to decide to press the left hand.","fun_headline_variants_meta":{"raw":{"variants":["Imagined tapping uses different timing rules than real tapping","Motor imagery and execution: distinct timing factors revealed","In imagined tasks, reaction times improve and track unpredictable cues","Imagined finger movements respond to different cues than executed ones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1218,"prompt_tokens":893,"completion_tokens":325,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":259}},"tokens_in":509,"tokens_out":325,"duration_ms":4019,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:25:32.941320+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Motor Imagery group with the variable-event probabilities reversed (V2 at 74% and V3 at 26%) and include a control condition in which the Motor Execution group performs the real tap and then presses the left-hand key to end the trial; the predicted inversion of the V2-after-V2 and V3-after-V3 reaction-time pattern in imagery, or a shift in the execution pattern toward the imagery pattern in the control, would directly test whether the distinct-factors claim is real or an artifact of the response structure.","supporting_citations":[],"review_version":1}