{"id":"e9e1438a-3d07-4562-95e2-824f2949851b","arxiv_id":"2509.10552","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Per-trial beta-band EEG desynchronization in frontal-central electrodes distinguishes Pain from No-Pain trials and tracks subjective pain ratings in 59 healthy adults.","lead":"This study tested whether moment-to-moment brain wave changes in EEG can signal pain without relying on self-report. In 59 healthy adults receiving mild electrical stimulation, a beta-band desynchronization in frontal-central electrodes distinguished painful from non-painful trials and tracked with the intensity people reported.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'marker' claim is supported only by in-sample GLMM fits on the same data used for ROI selection, so the reported p-values and effect sizes are not yet evidence of predictive utility; out-of-sample validation is needed.","rationale":"The reader's weakest_assumption identifies the ROI selection issue; I agree that this is a substantive concern, but I frame the most load-bearing version around the predictive claim, since the paper's clinical relevance rests on generalizability. The three headline results have different robustness: the Pain vs No-Pain contrast is so large and the selection was on the grand average (roughly orthogonal to the condition contrast in a balanced design) that it may well survive; the VAS prediction, however, is a small in-sample slope (r ≈ 0.14 in Pain) that is exactly the kind of effect that selection and overfitting can inflate. The stepwise procedure additionally invalidates the nominal p-values for all retained terms. Therefore the paper's own caveat ('preliminary evidence') is appropriate; the verdict CONDITIONAL is correct, pending out-of-sample validation and artifact release. I recommend no change to the reader's verdict.","tokens_in":8971,"tokens_out":7486,"duration_ms":68925,"concrete_test":"Leave-one-participant-out cross-validation of the full pipeline: for each held-out participant, re-run the TF mask/ROI selection and stepwise GLMM selection on the other 58 participants, then predict the held-out participant's trial-level VAS from their ROI-averaged ERD; compute out-of-sample R² and compare against a condition-only null model (e.g., by permutation of the ERD values within participant). If the out-of-sample R² is not positive or is not significantly better than the null, the 'nonverbal marker' claim is not supported by this study.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II-E defines the Beta ROI [220–760 ms; 15.21–30 Hz] from a grand-average TF significance mask on the same 59 participants, and four ROIs were identified, but only Beta is reported. Section II-F then applies backward stepwise elimination to a full factorial GLMM, yielding Equation 1, and Section III-B reports a 'reverse' GLMM that predicts VAS from ERD; neither the ROI selection nor the stepwise selection is adjusted for, and the VAS 'prediction' is an in-sample fit. While the condition contrast (p=7e-16) may be large enough to survive selection bias because the selection was based on the grand average, the VAS slope (p=3.42e-7), the Condition x VAS interaction (p=0.002), and all demographic interactions are unprotected. The conclusion that ERD 'supports its utility as a nonverbal marker of pain' depends on generalizability, which is not demonstrated by any split-sample or cross-validation analysis. No code or data are supplied to assess this.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript analyzes high-density EEG from 59 healthy participants receiving electrical stimulation under Pain and No-Pain conditions. Per-trial time-frequency decomposition is used to define regions of interest for event-related desynchronization (ERD), and generalized linear mixed models (GLMMs) relate beta-band ERD in frontal-central electrodes to experimental condition, subjective VAS intensity ratings, age, and gender. The authors report a strong Condition effect on beta ERD (p=7.02e-16), a Condition-by-VAS interaction (p=0.002), demographic interactions, and a 'reverse' model in which ERD predicts VAS across participants. They conclude that trial-level beta ERD supports utility as a nonverbal marker of pain.","tokens_in":9221,"tokens_out":2707,"duration_ms":28151,"significance":"If the central claim were supported, the paper would make a useful contribution by moving from trial-averaged analyses to trial-level mixed models and by quantifying demographic modulation of pain-related oscillatory responses. The study is a secondary analysis of an existing dataset, and the manuscript explicitly frames the results as preliminary. The main strengths are the use of per-trial data, a data-driven time-frequency masking procedure, and a mixed-model framework that accounts for participant-level variability. However, the evidence for the 'marker' claim is currently limited by in-sample analyses: the ROI is selected from the same data used for inference, the reduced model is obtained by unadjusted stepwise selection, and the reverse 'prediction' is a regression fit to the same data with no out-of-sample validation. The paper also does not provide code or data, which limits reproducibility and independent assessment. The significance of the reported effects is therefore uncertain until these validity threats are addressed.","major_comments":[{"comment":"The Beta ROI [220-760 ms; 15.21-30 Hz] is defined from a grand-average time-frequency significance mask computed on the same 59 participants used for all subsequent GLMMs, and four ROIs were identified while only the Beta ROI is reported. This creates a selection effect: p-values for the VAS interaction (p=0.002), the gender interaction (p=0.02), and demographic terms are not protected against the ROI-selection process. The very large Condition main effect (p=7.02e-16) may plausibly survive selection, but the weaker, load-bearing effects cannot be interpreted at face value. The authors should either report all ROIs (with appropriate correction), use a split-sample or cross-validated selection procedure, or explicitly state that the reported p-values are conditional on a data-derived ROI and therefore exploratory.","section":"II-E, III-A"},{"comment":"The 'reverse models' purportedly predicting VAS from ERD are in-sample fits to the same data used to fit the model; the equations in Figure 3 (VAS = 62.15 + 0.044*ERD) are ordinary regression lines, not predictive validation. No cross-validation, held-out participant, or out-of-sample evaluation is reported. The statement that ERD 'predicted VAS ratings across participants' and 'supports its utility as a nonverbal marker of pain' is therefore not supported by the presented analysis. At minimum, the manuscript should reframe these results as descriptive associations and add a genuine out-of-sample or cross-validated prediction analysis before making the marker claim.","section":"III-B, Fig. 3"},{"comment":"The statistical model is reduced by backward stepwise elimination, and only the final retained model is presented. Stepwise selection without adjustment (e.g., bootstrap, model-averaging, or selection-aware inference) means that the p-values for retained terms, including Condition-by-VAS and demographic interactions, are not valid as reported. The manuscript should either justify the selection procedure statistically, provide selection-adjusted estimates, or present the full factorial model and its fit statistics in addition to the reduced model.","section":"II-F, Eq. (1)"},{"comment":"A Gamma GLMM with a log link requires a strictly positive outcome, but ERD/ERS values are percentage changes that are negative during desynchronization. The manuscript states that 'transformed ERD values' were modeled but never specifies the transformation. Without this detail, the reported effect directions, the regression equations in Figure 3, and the model interpretation cannot be evaluated. The authors must state the exact transformation and how it relates to the ERD values displayed in the figures.","section":"II-F"}],"minor_comments":[{"comment":"The sentence 'see it in 1' appears to be an incomplete reference; it should read 'see Fig. 1' or similar.","section":"II-E"},{"comment":"The labels 'Alpha1', 'Alpha2', 'Alpha-Beta', and 'Beta' are introduced in the masking procedure, but only the Beta ROI is analyzed or discussed. The authors should state whether analyses for the other ROIs were performed and, if so, where the results are reported, or clarify the preregistered/planned focus.","section":"II-E"},{"comment":"The manuscript inconsistently uses 'V AS' with spaces; this should be changed to 'VAS' throughout for readability.","section":"Throughout"},{"comment":"The ICA artifact rejection thresholds (ICLabel 'Brain' probability < 0.5, and correlation > 0.4 with auxiliary channels) are reported, but the number of rejected components per participant is not summarized; reporting this would help assess data-quality variability.","section":"II-D"},{"comment":"The reported p-values in Section III-B (e.g., Condition p<2e-16, three-way interaction p=2e-6) are presented without effect sizes or confidence intervals; adding these would improve interpretability, especially given the in-sample concerns raised above.","section":"III-B"},{"comment":"The model comparison criteria (AIC, BIC, likelihood ratio tests) are mentioned but no table or numerical values are provided; a supplementary table with fit statistics for candidate models would help readers assess the stepwise reduction.","section":"II-F"}],"recommendation":"major_revision","confidential_remarks":"The paper is a secondary analysis of an existing dataset with no code or data availability statement. The central predictive claim is currently supported only by in-sample fits after data-driven ROI selection and unadjusted stepwise model selection. These problems are potentially fixable by reframing the results as exploratory and adding out-of-sample validation or selection-adjusted inference, but they are substantive enough that a major revision is appropriate. If the authors cannot obtain independent data or implement cross-validation, the conclusions should be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this paper has a real result buried under an overclaimed headline. Per-trial beta ERD in frontal-central electrodes (F1, F2, Fz, FCz) distinguishes Pain from No-Pain trials with p=7e-16 in 59 participants. That effect is large and consistent with prior literature. The paper also does something useful: it moves beyond trial averaging to single-trial GLMMs and includes age/gender interactions. The authors are appropriately cautious in the abstract — they call it preliminary evidence and say future validation is needed. Credit for that.\n\nThe soft spots are all about inference, not the raw phenomenon. The Beta ROI was selected from a grand-average TF significance mask computed on the same 59 participants, and only Beta is reported even though four ROIs were identified. The stepwise backward elimination for the GLMM is unadjusted. The 'reverse model' that predicts VAS from ERD is just an in-sample GLMM fit — no split-sample, no cross-validation. The small correlations (r~0.14) and p=3.42e-7 are not evidence of predictive utility. The Gamma distribution with log link is chosen for 'strictly positive' ERD, but percentage-change ERD can be negative; the actual transform used is never specified. No code or data are provided.\n\nNone of this falsifies the core association. The condition contrast is so strong (p~1e-15) that selection bias is unlikely to explain it entirely. But the VAS interaction (p=0.002), the demographic interactions, and the reverse-model p-values are unprotected and could easily be inflated. The paper's conclusion that ERD 'supports its utility as a nonverbal marker of pain' is one step too strong given the purely in-sample evidence.\n\nWho gets value: pain neuroscientists and anyone working on EEG-based markers. It is a good reading-group case study for why in-sample prediction is not prediction. I would send it to peer review — the phenomenon is important and the per-trial design is a step forward — but I would expect major revision: out-of-sample validation, adjustment for ROI selection and stepwise selection, full reporting of all ROIs, and artifact release. Without those, the predictive claim should not be published as is.\n\nRecommendation: engage with it, but treat the predictive claim as unproven.","headline":"Per-trial beta ERD distinguishes pain from no-pain with a strong effect, but the paper's predictive claim rests on in-sample fits and data-dependent ROI selection, so it is not yet evidence of a predictive marker.","tokens_in":9741,"tokens_out":2861,"would_cite":false,"duration_ms":25231,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Per-trial beta-band EEG desynchronization distinguishes pain from no-pain and predicts intensity ratings.","keywords":["event-related desynchronization","pain marker","EEG","beta band","trial-level analysis","generalized linear mixed model","visual analogue scale","time-frequency analysis"],"falsifier":"Re-run the same GLMMs on a pre-specified beta ROI and a held-out participant split: if the Pain–No-Pain effect and the ERD–VAS slope shrink toward zero out-of-sample, the marker claim is not generalizable; if they persist, the circularity concern is weakened.","tokens_in":8754,"feed_emoji":"🧠","tokens_out":8199,"duration_ms":72085,"temperature":0.7,"pith_summary":"The paper seeks to establish that a single-trial EEG feature—$\\beta$-band event-related desynchronization (ERD) over frontal-central electrodes—can act as a nonverbal, intensity-sensitive marker of pain. In 59 healthy adults receiving calibrated electrical stimulation, a data-driven $\\beta$ ROI [220–760 ms; 15.21–30 Hz] separated Pain from No-Pain trials in a generalized linear mixed model (main effect $p = 7.02\\times 10^{-16}$). ERD also scaled with subjective VAS ratings (Condition × VAS interaction $p = 0.002$), and reverse models showed ERD predicting VAS across participants ($p = 3.42\\times 10^{-7}$), with age and gender moderating the coupling. This matters because it points toward EEG-based pain monitoring for patients who cannot report pain, and toward quantitative endpoints for analgesic testing.","feed_headline":"Trial-level beta brain waves flag pain","feed_subtitle":"Frontal-central beta ERD separated painful from non-painful shocks and tracked intensity in 59 adults.","key_machinery":"The load-bearing object is the Beta ROI, a rectangular time–frequency patch defined by a data-driven masking procedure: pointwise t-tests of post-stimulus power against a pre-stimulus baseline distribution, family-wise error corrected across the grand-average time–frequency matrix, yielding the window [220–760 ms; 15.21–30 Hz] averaged over F1, F2, Fz, and FCz. ERD is expressed as percentage change relative to baseline power, and single-trial values are extracted by averaging within the ROI. Statistical inference is carried by Gamma generalized linear mixed models with log link and participant-level random intercepts, which model condition, VAS, age, gender, and interactions, plus reverse models that predict VAS from ERD. This design turns trial-level variability into the unit of analysis, which is the paper's main methodological departure from trial-averaged pain EEG studies.","core_discovery":"The central claim is that per-trial beta-band ERD, measured as baseline-relative power decrease in electrodes F1, F2, Fz, and FCz, carries both categorical and intensity information about pain perception. Pain trials produced stronger desynchronization than No-Pain trials, and the effect was robust across model specifications. The relationship with subjective intensity was significant but directionally surprising: within the Pain condition, higher VAS ratings were associated with less negative ERD, i.e., beta power stayed closer to baseline when participants reported higher pain. The same feature predicted VAS ratings in reverse models, and demographic variables—age and gender—moderated the ERD–VAS coupling. The paper frames this as preliminary evidence that trial-level EEG oscillations can serve as reliable indicators of pain and support individualized, report-free pain monitoring.","pith_inferences":["If the ROI-selection step is circular in effect, the headline p-values are optimistic; a pre-registered split-sample reanalysis with a fixed ROI is the direct way to estimate the true generalization error.","A specificity test comparing painful with non-painful but equally salient somatosensory stimulation would show whether the beta ERD signature encodes pain perception or general stimulus salience.","The positive ERD–VAS slope in Pain trials could be exploited therapeutically: if frontal beta maintenance reduces pain, beta-enhancing or alpha-entraining visual stimulation might modulate the same circuit, an extension the paper's discussion only gestures at."],"forward_implications":["If the finding is correct, frontal-central beta ERD becomes a continuous, nonverbal readout of pain intensity rather than only a Pain/No-Pain classifier.","The marker could support pain monitoring in non-communicative patients and provide an objective endpoint for analgesic or neuromodulatory trials.","Because age and gender change the ERD–VAS coupling, pain-decoding models will need demographic covariates rather than a single population-level slope.","The inverse ERD–VAS direction suggests that beta activity near baseline at high pain may reflect compensatory or top-down processes, a hypothesis that follow-up attention-manipulation studies can test."],"supporting_citations":[{"why":"Supplies the parent dataset of 59 healthy participants with Pain/No-Pain electrical stimulation EEG and VAS ratings reanalyzed here.","marker":"[16]"},{"why":"Establishes brain rhythms as candidate pain markers and motivates moving beyond trial averaging to trial-level oscillatory decoding.","marker":"[1]"},{"why":"Previous evidence that oscillations differentially encode noxious stimulus intensity versus perceived pain, framing the ERD–VAS coupling result.","marker":"[3]"},{"why":"Prior work on neural indicators of perceptual variability of pain that the per-trial GLMM approach extends.","marker":"[4]"},{"why":"Provides the complex Morlet wavelet time-frequency convolution method used to compute single-trial ERD.","marker":"[20]"},{"why":"Supplies the lme4 generalized linear mixed-model machinery used for both the condition and VAS models.","marker":"[21]"}],"fun_headline_variants":["Per-trial beta desync separates pain from no-pain","Beta dip on each shock tracks pain strength","EEG beta power drop per trial predicts pain","No-averaging beta ERD reveals pain per trial","Trial-wise beta ERD: a nonverbal pain marker"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The strongest numbers rest on defining the beta region of interest from a significance mask on the same 59-participant dataset that is then used to test the condition and intensity effects, so the analysis assumes this does not inflate significance.","fun_headline_variants_meta":{"raw":{"variants":["Per-trial beta desync separates pain from no-pain","Beta dip on each shock tracks pain strength","EEG beta power drop per trial predicts pain","No-averaging beta ERD reveals pain per trial","Trial-wise beta ERD: a nonverbal pain marker"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000759,"raw_usage":{"total_tokens":3365,"prompt_tokens":930,"completion_tokens":2435,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":2358}},"tokens_in":546,"tokens_out":2435,"duration_ms":16796,"temperature":1.0,"reasoning_tokens":2358,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:12:50.325272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same GLMMs on a pre-specified beta ROI and a held-out participant split: if the Pain–No-Pain effect and the ERD–VAS slope shrink toward zero out-of-sample, the marker claim is not generalizable; if they persist, the circularity concern is weakened.","supporting_citations":[{"cited_title":"The placebo hypoalgesic response is reduced in healthy older adults showing a decline in executive functioning,","cited_arxiv_id":null,"evidence_quote":"Supplies the parent dataset of 59 healthy participants with Pain/No-Pain electrical stimulation EEG and VAS ratings reanalyzed here."},{"cited_title":"Brain rhythms of pain,","cited_arxiv_id":null,"evidence_quote":"Establishes brain rhythms as candidate pain markers and motivates moving beyond trial averaging to trial-level oscillatory decoding."},{"cited_title":"Brain oscillations differentially encode noxious stimulus intensity and pain intensity,","cited_arxiv_id":null,"evidence_quote":"Previous evidence that oscillations differentially encode noxious stimulus intensity versus perceived pain, framing the ERD–VAS coupling result."},{"cited_title":"Neural indicators of perceptual variability of pain across species,","cited_arxiv_id":null,"evidence_quote":"Prior work on neural indicators of perceptual variability of pain that the per-trial GLMM approach extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the complex Morlet wavelet time-frequency convolution method used to compute single-trial ERD."},{"cited_title":"Fitting linear mixed- effects models using lme4,","cited_arxiv_id":null,"evidence_quote":"Supplies the lme4 generalized linear mixed-model machinery used for both the condition and VAS models."}],"review_version":1}