Pith. sign in

REVIEW 1 major objections 3 minor 16 references

Frozen EEG embeddings miss long-range temporal dynamics, while preserving static spectral shape.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 01:31 UTC pith:OMO7SZOF

load-bearing objection Careful, self-critical EEG-FM audit whose central spectral–temporal dissociation is undercut by R² values that are not comparable across cohorts or targets. the 1 major comments →

arxiv 2607.24834 v2 pith:OMO7SZOF submitted 2026-07-23 q-bio.NC cs.AIcs.ETcs.LG

Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations

classification q-bio.NC cs.AIcs.ETcs.LG
keywords EEG foundation modelsdetrended fluctuation analysislong-range temporal correlationscross-population transferscale-free dynamicsaperiodic exponentalpha enveloperecording-site leakage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Frozen EEG foundation-model representations are commonly used as general-purpose brain-signal features, but it is unclear whether they preserve the temporal order of the signal. This paper tests whether five such models—REVE, LaBraM, BENDR, CBraMod, and BIOT—can decode a subject's long-range temporal correlations, measured as the detrended-fluctuation-analysis exponent of the alpha-band amplitude envelope, in two independent cohorts. With a single fixed readout, CBraMod and BIOT predicted the aperiodic 1/f exponent in both cohorts, but DFA decoding was cohort-dependent: only BIOT was positive in CAUEEG, and none of the five replicated in BrainLat. The claimed discovery is a model-specific spectral–temporal dissociation: static spectral shape is partly preserved, but temporal-scaling structure is not reproducibly encoded. If correct, it bounds what frozen embeddings can support for LRTC-based clinical biomarkers and motivates an LRTC-aware pretraining objective.

Core claim

The paper's central claim is that frozen EEG foundation models show a dissociation between what they preserve from the static frequency domain and what they preserve from the temporal order of the signal. Using a common 240 s estimator, artifact masking, and one fixed PCA-ridge readout, CBraMod and BIOT decoded the aperiodic exponent in both cohorts (CAUEEG R^2 = 0.459 and 0.606; BrainLat R^2 = 0.652 and 0.757), whereas alpha-envelope DFA decoding was positive only in CAUEEG for BIOT (R^2 = 0.232) and negative in BrainLat for all five models. The dissociation is model-specific because the two models that succeeded on the spectral target are the same two that failed to reproduce the temporal

What carries the argument

The central mechanism is a controlled comparison of two decoding targets from the same frozen embeddings: the DFA exponent, a dimensionless measure of long-range temporal correlations computed from the detrended fluctuation slope of the alpha-envelope over 2–23.8 s windows, and the aperiodic 1/f exponent from fixed-mode spectral parameterization. The probe is a fixed PCA-ridge regression readout evaluated with nested cross-validation, with diagnostic controls that shuffle pre-pool token order and residualize the aperiodic association. The dissociation itself—same embeddings, same readout, one target decodable in both cohorts and the other not—is the load-bearing result, and it is used to arg

Load-bearing premise

The load-bearing assumption is that the DFA target is as reliable as the aperiodic target, so the large gap in R-squared reflects what the embeddings encode rather than a difference in how much measurement noise the two targets carry; the paper measures half-target repeatability only for the shorter scale range and gives no reliability estimate for the aperiodic target.

What would settle it

Measure the reliability of the full 2–23.8 s DFA exponent (e.g., split-half on longer recordings or test-retest) and compare it with the reliability of the aperiodic exponent. If the DFA target is substantially less reliable, the lower and cohort-dependent DFA R-squared would be expected even if the embeddings carried identical information. Conversely, observing the CAUEEG DFA decoding replicate in a third independent cohort with documented high target reliability would settle the dissociation in the paper's favor.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the dissociation is correct, frozen-embedding pipelines should not be assumed to support LRTC-based biomarkers such as altered alpha-envelope scaling in dementia; each temporal feature needs direct validation.
  • The result extends the known spectral bias of reconstruction-based EEG models to a temporal feature: preserving 1/f shape does not imply preserving the order-dependent scaling exponent.
  • The proposed LRTC-aware auxiliary objective—a differentiable scale-freeness and exponent-fidelity penalty added to pretraining—is a concrete, testable intervention that would either support or refute the claim that current objectives under-reward temporal ordering.
  • Cross-cohort failure of DFA decoding, while aperiodic decoding replicates, implies that cohort, site, or population factors interact with what these embeddings encode, so cross-population generalization needs explicit auditing.
  • A null transfer result under source-label permutation means that favorable target-bootstrap intervals are not sufficient evidence for zero-shot transfer of DFA-based features.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit: if temporal ordering is the lost axis, other order-dependent features—phase-amplitude coupling, recurrence measures, or state-transition dynamics—should also fail to decode reproducibly from the same embeddings; testing them would map the boundary of the dissociation.
  • The cohort dependence could be tested by applying the same probe to a third independent cohort; a reproduction of the CAUEEG positive DFA result in another cohort with high target reliability would strengthen the dissociation, while a null in a well-powered cohort would suggest the CAUEEG finding was site-specific.
  • The paper proposes an intervention but does not evaluate it; a synthetic benchmark with fractional Gaussian noise of known Hurst exponent, where the exponent is the entire signal, would isolate whether the failure is about temporal order per se or about the interaction of LRTC with spectral structure.
  • One could also fine-tune only the top layers of a frozen model on the DFA target and observe whether decoding improves; if it does, the information is present but not linearly accessible, which would reframe the result as a readout limitation rather than an encoding one.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 3 minor

Summary. The paper tests whether frozen embeddings from five EEG foundation models (REVE, LaBraM, BENDR, CBraMod, BIOT) support decoding of the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope, a measure of long-range temporal correlations (LRTC), in two independent dementia cohorts (CAUEEG, N=764; BrainLat, N=79). Using one fixed nested-CV PCA-ridge readout, the authors report that BIOT and CBraMod decode DFA in CAUEEG (R^2=0.232 and 0.121, respectively) but not in BrainLat, where all five models have negative point estimates. In contrast, CBraMod and BIOT decode the aperiodic 1/f exponent in both cohorts (R^2=0.459–0.757). The paper interprets this as a model-specific spectral–temporal dissociation, while carefully caveating that negative decoding is readout-specific, transfer is unconfirmed, and several controls are descriptive rather than confirmatory.

Significance. If the dissociation holds, the result is useful: it suggests that current frozen EEG-FM representations preserve static spectral shape for some models but do not reproducibly encode temporal-scaling structure, bounding their use for LRTC-based analyses. The paper is unusually transparent about its adaptive analysis path, limits, and negative controls; it ships code and audit metadata, and it explicitly disclaims overinterpretation (e.g., not claiming representational absence, not treating rank-order controls as LRTC-specific, not treating ComBat as causal). These are real strengths. However, the central claim currently rests on comparing R^2 values across targets and cohorts without accounting for differences in target variance, which is a load-bearing gap.

major comments (1)
  1. [§4.2, Tables 2/3] The dissociation claim compares DFA R^2 with aperiodic R^2 across cohorts, but R^2 = 1 − MSE/Var(target) is not comparable across targets with different between-subject variance. Table 2 shows BrainLat DFA SD is roughly half that of CAUEEG (0.048–0.057 vs 0.100–0.109), i.e., ~4× less variance. Concretely, BIOT's BrainLat DFA R^2 = −0.06 implies MSE ≈ 0.0029 (using SD ≈ 0.052), whereas its CAUEEG R^2 = 0.232 implies MSE ≈ 0.0085; absolute error is smaller in BrainLat, yet R^2 is negative. The aperiodic target's variance is not reported. Thus the spectral–temporal contrast could be a target-variance artifact rather than a representational difference. Please report MSE, target SD, and a variance-aware comparison (e.g., standardized MSE or R^2 on variance-matched subsamples) to support the central claim.
minor comments (3)
  1. [Table 4 note] The note 'REVE uses finer tokens (N=396); the other models use 60 four-second epoch embeddings (N=199)' is ambiguous: does N refer to subjects or token/epoch counts? Clarify, since the CAUEEG DFA analysis uses N=764 elsewhere.
  2. [Throughout] The manuscript inconsistently types the exponent as 'DFA' and 'DF A' (e.g., abstract vs. Section 3). Please unify.
  3. [§4.1] The half-target repeatability r=0.824 is reported only for the 2–11.8 s estimand; consider also reporting split-half reliability for the full 2–23.8 s estimand if computationally feasible, given its direct relevance to the target's measurement quality.

Circularity Check

0 steps flagged

No significant circularity; the paper is an empirical encoding-probe evaluation, not a derivation, and its conclusions do not reduce to fitted parameters or self-citations.

full rationale

The paper's central claim—that CBraMod and BIOT decode the aperiodic exponent in both cohorts while alpha-envelope DFA decoding is cohort-dependent—is an empirical result from out-of-fold ridge regressions on frozen embeddings. The targets (DFA-v2 and fixed-mode aperiodic slope) are defined independently of the embeddings and of each other; the readout is fixed and evaluated by nested CV, so the R^2 values are genuine held-out predictions rather than fitted parameters renamed as predictions. The target computations (FIR 8-13 Hz filtering, DFA over 2-23.8 s, QC, fixed-mode 1-40 Hz spectral parameterization) are standard and are not defined in terms of the embeddings or of each other. The aperiodic-residualization control directly tests target redundancy rather than assuming it. Self-citations (e.g., Zare 2026) are contextual and not load-bearing; none is invoked to force the dissociation. The R^2=0.10 reference line is explicitly admitted to be a post-revision descriptive aid, not a prespecified threshold, and the paper does not use it as a decision boundary. The main weaknesses—half-target repeatability not being a ceiling, no reliability reported for the aperiodic target, and possible differences in target variance affecting cross-target/cohort comparability of R^2—are validity or measurement concerns, not circular reasoning. The paper itself flags its controls as descriptive and its transfer as unconfirmed, reinforcing that no result is being justified by its own definition. Score 1 reflects only minor post-hoc/adaptive elements and self-citation presence, none of which is load-bearing.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The central claims rest on the DFA estimand, cohort harmonization, pretraining non-overlap, and estimator-bias assumptions listed above. No new physical or model entities are introduced; the proposed LRTC-aware loss is a method suggestion, not an invented entity.

free parameters (4)
  • DFA scale lower bound = 2 s
    Chosen as the fixed lower bound after simulations showed mean null exponent 0.538 with residual bias +0.038 at 2 s, versus 0.600 at 0.5 s. This hand-chosen bound defines the target.
  • QC eligibility thresholds = >=15 finite channels; median retained fraction >=0.80; median log-log fit R^2 >=0.95
    Author-selected inclusion criteria determine which recordings (764/79) enter every analysis; these thresholds are not derived from an external benchmark.
  • PCA component cap and alpha band = 50 components; 8-13 Hz
    The PCA cap bounds the linear readout and affects achievable R^2. The alpha band is standard, but it is a fixed choice rather than a derived quantity.
  • R^2=0.10 descriptive reference = 0.10
    Adopted during the post-v1 revision after results were known; the paper states it is not a prespecified threshold and is not used for inference, but it shapes interpretation.
axioms (5)
  • domain assumption DFA of the alpha-band amplitude envelope is a valid operationalization of long-range temporal correlations.
    The paper follows Hardstone et al. (2012) and Linkenkaer-Hansen et al. (2001), treating DFA as the LRTC target while explicitly separating DFA from criticality claims.
  • domain assumption Harmonization to a common 19-channel montage, 200 Hz, common-average reference, and first 240 s preserves cross-cohort comparability of both targets.
    Cohorts differ in hardware, montage, and recording length; the paper notes DFA is not invariant to re-referencing or channel mixing, so comparability is assumed.
  • domain assumption The pretraining corpora of LaBraM, BENDR, CBraMod, and BIOT do not contain the test cohorts.
    Only REVE's documented corpus was checked; Section 7 says overlap cannot be excluded for the other models. Interpretation of decoding as representation transfer rather than memorization depends on this.
  • domain assumption The residual DFA estimator bias (+0.038) and the one-decade scale range do not differentially affect models or cohorts.
    The bias and scale-range limitations are reported but not corrected; the paper assumes they do not drive the decoding contrast.
  • standard math Source-label permutation tests are valid under the null.
    Permutation is standard and the paper refits the full model, but validity assumes exchangeability of source labels under the null.

pith-pipeline@v1.3.0-alltime-deepseek · 14945 in / 13921 out tokens · 134379 ms · 2026-08-04T01:31:52.384939+00:00 · methodology

0 comments
read the original abstract

Objective. We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, measured as the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope. Approach. REVE, LaBraM, BENDR, CBraMod, and BIOT were evaluated in CAUEEG and BrainLat. A common 240 s estimator used 8-13 Hz filtering, DFA over 2-23.8 s, artifact masking, and quality control. One fixed nested-cross-validation readout predicted DFA and a fixed-mode aperiodic exponent. Controls tested pre-pool order sensitivity and aperiodic residualization. Results. CAUEEG included 764 recordings and BrainLat 79. BIOT decoded DFA in CAUEEG (R-squared = 0.232; conditional subject-bootstrap 95 percent interval, 0.121-0.310), and CBraMod was positive but imprecise (R-squared = 0.121; 0.003-0.214). Neither replicated in BrainLat, where all five point estimates were negative. In contrast, CBraMod and BIOT decoded the aperiodic exponent in both cohorts (R-squared = 0.459-0.757). BIOT remained positive after removal of the measured linear aperiodic association in matched CAUEEG data (R-squared = 0.240). The post-hoc order control was batch- and configuration-sensitive. Because chronological EEG epochs are not exchangeable, it was descriptive, not an LRTC-specific test. No revised DFA transfer direction passed source-label permutation testing. Cohort membership was near-ceiling decodable from all five embeddings, but this is not a pure site effect. Significance. CBraMod and BIOT show a replicated, model-specific spectral-temporal dissociation: aperiodic decoding is present in both cohorts, whereas alpha-envelope DFA decoding is cohort-dependent. These findings bound the evaluated readouts; they do not establish representational absence or an architectural cause. Transfer and clinical associations remain exploratory.

Figures

Figures reproduced from arXiv: 2607.24834 by Marzieh Zare.

Figure 1
Figure 1. Figure 1: The dissociation, both cohorts. Bars show the best cross-validated [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 1
Figure 1. Figure 1: Distributional quality control for the revised target among eligible recordings. Curves [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Predicted vs. true alpha-envelope DFAfull exponent on CAUEEG (N = 770, 5-fold CV, ridge probe). Left: the classical feature tracks the identity line. Right: REVE’s predictions collapse to a narrow band near the sample mean irrespective of the true value, visually confirming that R2 ≈ 0 reflects a genuine failure to recover subject-level LRTC structure rather than an artifact of the target’s narrow variance… view at source ↗
Figure 2
Figure 2. Figure 2: Fixed-readout encoding results. Points are out-of-fold [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Pre-pool order control (CAUEEG, DFA, 5 seeds). REVE and CBraMod show [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 3
Figure 3. Figure 3: Observed and out-of-fold predicted DFA-v2 values. Dashed lines are the identity. The [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Cross-population DFA transfer (point AUROC, resample 95% CI, raw permutation [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 4
Figure 4. Figure 4: Post-hoc pre-pool order control against DFA v2. Grey points are the configuration-specific [PITH_FULL_IMAGE:figures/full_fig_p012_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

16 extracted references · 2 canonical work pages · 1 internal anchor

  1. [1]

    A spectral audit framework reveals task-dependent aperiodic reliance across eeg and ecg deep learning.arXiv preprint arXiv:2606.08583,

    Jasmeet Singh Bindra, Siddharth Panwar, and Shubhajit Roy Chowdhury. A spectral audit framework reveals task-dependent aperiodic reliance across eeg and ecg deep learning.arXiv preprint arXiv:2606.08583,

  2. [5]

    Eeg-bench: A benchmark for eeg foundation models in clinical applications.arXiv preprint arXiv:2512.08959,

    Ard Kastrati, Josua B¨ urki, Jonas Lauer, Cheng Xuan, Raffaele Iaquinto, and Roger Wattenhofer. Eeg-bench: A benchmark for eeg foundation models in clinical applications.arXiv preprint arXiv:2512.08959,

  3. [9]

    Beyond accuracy: Ro- bustness, interpretability and expressiveness of EEG foundation models.arXiv preprint arXiv:2605.17562,

    Urban ˇSirca, Maryam Alimardani, Stefanos Zafeiriou, and Konstantinos Barmpas. Beyond accuracy: Ro- bustness, interpretability and expressiveness of EEG foundation models.arXiv preprint arXiv:2605.17562,

  4. [10]

    Pretrained, frozen, still leaking: Auditing cross-encoder attribute transfer in eeg foundation models.arXiv preprint arXiv:2606.09189,

    Jianwei Tai. Pretrained, frozen, still leaking: Auditing cross-encoder attribute transfer in eeg foundation models.arXiv preprint arXiv:2606.09189,

  5. [11]

    What do eeg foundation models capture from human brain signals?arXiv preprint arXiv:2605.11410,

    Ling Tang, Qian Chen, Jilin Mei, Houshi Xu, Quanshi Zhang, Jing Shao, Na Zou, Xia Hu, and Dongrui Liu. What do eeg foundation models capture from human brain signals?arXiv preprint arXiv:2605.11410,

  6. [12]

    Baker, Yu Wu, Anand D

    Ye Tao, Bradley T. Baker, Yu Wu, Anand D. Sarwate, Sandeep Panta, Sergey Plis, and Vince D. Calhoun. Batch effects in brain foundation model embeddings.arXiv preprint arXiv:2604.14441,

  7. [14]

    Chaoqi Yang, M

    arXiv:2412.07236. Chaoqi Yang, M. Brandon Westover, and Jimeng Sun. Biot: Biosignal transformer for cross-data learning in the wild. InAdvances in Neural Information Processing Systems (NeurIPS),

  8. [15]

    Stress-testing EEG foundation models for clinical decoding: Dataset identity and targeted negative controls.arXiv preprint arXiv:2607.24519,

    Marzieh Zare. Stress-testing EEG foundation models for clinical decoding: Dataset identity and targeted negative controls.arXiv preprint arXiv:2607.24519,

  9. [16]

    What EEG Foundation Models Encode: Dataset Identity and a Negative-Control Suite for Clinical Benchmarks

    doi: 10.48550/arXiv.2607.24519. URLhttps: //arxiv.org/abs/2607.24519. Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. Are transformers effective for time series forecasting? InProceedings of the AAAI Conference on Artificial Intelligence,

  10. [2006]

    Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Haiteng Jiang, Shijian Li, Tao Li, and Gang Pan

    doi: 10.1186/1471-2105-7-91. Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Haiteng Jiang, Shijian Li, Tao Li, and Gang Pan. Cbramod: A criss-cross brain foundation model for eeg decoding. InInternational Conference on Learning Representations (ICLR),

  11. [2017]

    doi: 10.1177/1948550617697177. William Lehn-Schiøler, Magnus Ruud Kjær, Rahul Thapa, Magnus Guldberg Pedersen, Anton Mosquera Storgaard, Nick Williams, Radu Gatej, Tue Lehn-Schiøler, Andreas Brink-Kjær, Sadasivan Puthussery- pady, S´ andor Beniczky, James Zou, and Lars Kai Hansen. Mechanistic interpretability of EEG foundation models via sparse autoencode...

  12. [2018]

    doi: 10.1016/j.neuroimage.2017. 11.024. Richard Hardstone, Simon-Shlomo Poil, Giuseppina Schiavone, Rick Jansen, Vadim V. Nikulin, Huibert D. Mansvelder, and Klaus Linkenkaer-Hansen. Detrended fluctuation analysis: A scale-free view on neuronal oscillations.Frontiers in Physiology, 3:450,

  13. [2023]

    CAUEEG: Chung-Ang University Hospital EEG dataset. Aditya Kommineni, Emily Zhou, Kleanthis Avramidis, Simon Bock Segaard, Jeppe Roden M¨ unster, An- dreas Peter Juhl Hansen, Takfarinas Medani, Tiantian Feng, Richard Leahy, and Shrikanth Narayanan. Aperiodic and low-frequency spectral bias in reconstruction-based eeg foundation models.arXiv preprint arXiv:...

  14. [2024]

    Yiru Jiao, Sander van Cranenburgh, Simeon Calvert, and Hans van Lint

    spotlight; arXiv:2405.18765. Yiru Jiao, Sander van Cranenburgh, Simeon Calvert, and Hans van Lint. Structure-preserving contrastive learning for spatial time series.arXiv preprint arXiv:2502.06380,

  15. [2025]

    Jean-Philippe Fortin, Nicholas Cullen, Yvette I

    arXiv:2510.21585. Jean-Philippe Fortin, Nicholas Cullen, Yvette I. Sheline, Warren D. Taylor, Irem Aselcioglu, Philip A. Cook, Phil Adams, Crystal Cooper, Maurizio Fava, Patrick J. McGrath, Melvin McInnis, Mary L. Phillips, Madhukar H. Trivedi, Myrna M. Weissman, and Russell T. Shinohara. Harmonization of cortical thickness measurements across scanners an...

  16. [2026]

    The identity trap in eeg foundation models: A diagnostic audit.arXiv preprint arXiv:2606.06647,

    Jun-You Lin, Ying Choon Wu, and Tzyy-Ping Jung. The identity trap in eeg foundation models: A diagnostic audit.arXiv preprint arXiv:2606.06647,