Pith. sign in

REVIEW 38 references

Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read FastText and GPT-2 embeddings linearly predict sEEG high-gamma responses during word reading with high correlation, but Wav2Vec 2.0 performs poorly in two participants.

desk verdict Solid feasibility study on a public sEEG dataset, but the abstract overclaims relative to the authors' own Table 1 and the missing null baseline leaves the main result ambiguous. read the letter →

arxiv 2505.14074 v2 pith:NOAAXYQZ submitted 2025-05-20 cs.HC cs.SDeess.AS

classification cs.HCcs.SDeess.AS
keywords speechactivityembeddingshigh-gammalanguageneuralproductioncorrelation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Deep learning models trained on text and audio contain internal representations of words and speech. This paper asks whether those representations can predict brain activity recorded while people speak. The authors use a public dataset of 10 people with implanted depth electrodes who read 100 Dutch words aloud. From the electrode signals they extract high-gamma power, a frequency band linked to speech production. For each word, they compute three kinds of embeddings: FastText word vectors, averaged GPT-2 token vectors, and Wav2Vec 2.0 audio embeddings from the participant's own voice.

For each participant and embedding type, they train an ElasticNet linear model on 99 words and test it on the remaining word, repeating until every word has been held out once. The model predicts the full pattern of high-gamma activity across electrodes and time windows. Reconstruction quality is measured with Pearson correlation and R2.

The text-based embeddings reach high correlations for most participants, with subject S04 near 0.99. The audio-based Wav2Vec model is unstable, giving near-zero or negative R2 for subjects S07 and S09 and a mean PCC of only 0.25 for S09. The abstract's claim of correlations from 0.79 to 0.99 for all participants is inconsistent with the paper's own Table 1. The evaluation also has no baseline comparing random embeddings or a simple mean-response model, so part of the high correlation may reflect predictable word-evoked activity rather than specific information in the embeddings.

Extended reading notes

Core claim

The abstract states that high-gamma activity can be effectively reconstructed using large language and speech model embeddings in all study participants, generating Pearson's correlation coefficients ranging from 0.79 to 0.99. If correct, this means pre-trained word or speech embeddings contain enough information to linearly predict the spatial and temporal pattern of high-gamma sEEG activity during word reading.

Load-bearing premise

The leave-one-out evaluation assumes that 100 unique, non-repeated words are independent samples and that a linear map from a static word embedding to the full spatiotemporal high-gamma response generalizes across words. Because each word appears once and the input is word-level, the high correlation could be driven by common word-evoked response patterns shared by all words, not by the specific embedding content. This is asserted in Section 2.6 without a null baseline.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper contributes an empirical evaluation rather than a theoretical derivation, so the ledger mainly records hand-chosen preprocessing and per-subject fitted regression weights. The largest unacknowledged cost is the absence of a baseline model that would show the embeddings, rather than generic word-evoked activity, are responsible for the high correlations.

free parameters (4)
  • ElasticNet regularization parameters lambda1 and lambda2 = not reported
    Chosen by hand and not reported; these control sparsity and directly affect the reconstructed signals.
  • PCA variance retained for Wav2Vec embeddings = 90%
    Hand-chosen threshold in Section 2.4; the paper does not state whether PCA is fit on training words only, so test information could leak into the transformation.
  • Per-subject ElasticNet weight matrices W = fit per subject per embedding type
    The central reconstruction is produced by these fitted matrices; weights are not reported, making the result a fitted black box.
  • Selected embedding layers = GPT-2 penultimate layer, Wav2Vec last layer, FastText static 300d
    Manual choices in Sections 2.3 and 2.4; performance may depend on which layer is used.
assumptions (5)
  • domain assumption High-gamma band 70-170 Hz reflects speech and language processing
    Invoked in Section 2.2 to define the target neural feature; the paper relies on prior evidence that this band encodes speech-language processing.
  • domain assumption Sparse, clinically determined electrode placement is adequate for participant-level claims
    Section 2.1 describes electrode placement determined by clinical needs; the paper assumes the sparse, variable coverage supports the reported reconstruction accuracy.
  • domain assumption Linear (ElasticNet) mapping between embeddings and high-gamma features is sufficient
    Section 2.5 uses ElasticNet; the paper assumes linearity with no nonlinear baselines to justify this choice.
  • domain assumption Leave-one-out across 100 unique words gives unbiased generalization
    Section 2.6 assumes each word is an independent sample; no repeated words means word-level variance cannot be separated from trial noise.
  • domain assumption Pre-trained embeddings transfer to Dutch speech production
    Sections 2.3 and 2.4 assume FastText/GPT-2 trained on English and internet text and XLS-R trained on multilingual audio transfer to Dutch production without adaptation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings." pith.science (2026). https://pith.science/paper/NOAAXYQZ

@misc{pith2026250514074,
  author       = {Pith},
  title        = {Pith review of: Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NOAAXYQZ}},
  note         = {Machine review of arXiv:2505.14074}
}
read the original abstract

Understanding how neural activity encodes speech and language production is a fundamental challenge in neuroscience and artificial intelligence. This study investigates whether embeddings from large-scale, self-supervised language and speech models can effectively reconstruct high-gamma neural activity characteristics, key indicators of cortical processing, recorded during speech production. We leverage pre-trained embeddings from deep learning models trained on linguistic and acoustic data to represent high-level speech features and map them onto these high-gamma signals. We analyze the extent to which these embeddings preserve the spatio-temporal dynamics of brain activity. Reconstructed neural signals are evaluated against high-gamma ground-truth activity using correlation metrics and signal reconstruction quality assessments. The results indicate that high-gamma activity can be effectively reconstructed using large language and speech model embeddings in all study participants, generating Pearson's correlation coefficients ranging from 0.79 to 0.99.

Figures

Figures reproduced from arXiv: 2505.14074 by the authors.

Figure 1
Figure 1. Methodology overview. 2. Method [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. R 2 scores for the FastText embeddings. 01 02 03 04 05 06 07 08 09 Subjects 0.6 0.4 0.2 0.0 0.2 0.4 0.6 0.8 1.0 R² Score [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. R 2 scores for the GPT-2.0 embeddings. 3. Results 3.1. Performance of text-based embeddings Figs. 2 and 3 present the R 2 scores for the FastText and GPT￾2.0 models, respectively. Both models exhibit similarities in their predictive performance, with certain subjects consistently achieving high R 2 values. Notably, subject S04 achieved near￾perfect reconstruction, with a mean R 2 score of 0.99 for both FastText and … view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: R 2 scores for the Wav2Vec 2.0 embeddings. 0.0 - 0.1 0.1 - 0.2 0.2 - 0.3 0.3 - 0.4 0.4 - 0.5 0.5 - 0.6 0.6 - 0.7 0.7 - 0.8 0.8 - 0.9 0.9 - 1.0 R² 0 10 20 30 40 50 No. of Trials 0 1 0 1 1 2 3 8 24 52 0 1 0 1 2 1 3 7 24 53 0 1 1 1 2 4 4 10 21 43 GPT-2.0 FastText Wave [P…
Figure 5
Figure 5. Figure 5: Average distribution of trials based on R 2 scores. 3.2. Performance of Wav2Vec 2.0 embeddings [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

38 extracted references · 32 canonical work pages

  1. [1]

    Introduction Speech neuroprostheses, also known as speech brain-computer interfaces (BCI), are a class of assistive technology aimed at restoring or increasing communication abilities for individu- als with severe speech and motor impairments. These de- vices decode neural activity associated with speech production and translate it into text [1, 2, 3, 4, ...

  2. [2]

    1 presents a block diagram summarizing our methodology for this work

    Method Fig. 1 presents a block diagram summarizing our methodology for this work. We used high-gamma features extracted from sEEG signals as the target neural features that were predicted from the embeddings computed by large, self-supervised audio and text models. These recordings were obtained from 10 par- ticipants implanted with depth electrodes while...

  3. [3]

    Performance of text-based embeddings Figs

    Results 3.1. Performance of text-based embeddings Figs. 2 and 3 present theR 2 scores for the FastText and GPT- 2.0 models, respectively. Both models exhibit similarities in their predictive performance, with certain subjects consistently achieving highR 2 values. Notably, subject S04 achieved near- perfect reconstruction, with a meanR 2 score of 0.99 for...

  4. [4]

    Our experi- ments provide evidence that the neural activity associated with speech production can be effectively modeled using machine learning and deep learning

    Conclusions We explored the potential for speech neuroprostheses by leveraging embeddings from the pre-trained large-scale, self- supervised language and audio models to reconstruct the neu- ral activity recorded during speech production. Our experi- ments provide evidence that the neural activity associated with speech production can be effectively model...

  5. [5]

    We are grateful to Dr

    Acknowledgments This work was supported by the grant PID2022-141378OB- C22 funded by MICIU/AEI/10.13039/501100011033 and ERDF/EU. We are grateful to Dr. Christian Herff from Maas- tricht University for generously sharing the sEEG dataset with us

  6. [6]

    Brain-to-text: decoding spoken phrases from phone representations in the brain,

    C. Herff, D. Heger, A. De Pesters, D. Telaar, P. Brunner, G. Schalk, and T. Schultz, “Brain-to-text: decoding spoken phrases from phone representations in the brain,”Frontiers in neu- roscience, vol. 9, p. 217, 2015

  7. [7]

    Neuroprosthesis for decoding speech in a paralyzed person with anarthria,

    D. A. Moses, M. K. Leonard, J. G. Makin, and E. F. Chang, “Neuroprosthesis for decoding speech in a paralyzed person with anarthria,”New England Journal of Medicine, vol. 385, no. 3, pp. 217–227, 2021

  8. [8]

    Generalizable spelling using a speech neuroprosthesis in an individual with severe limb and vocal paral- ysis,

    S. L. Metzger, J. R. Liu, D. A. Moses, M. E. Dougherty, M. P. Seaton, K. T. Littlejohn, J. Chartier, G. K. Anumanchipalli, A. Tu- Chan, K. Gangulyet al., “Generalizable spelling using a speech neuroprosthesis in an individual with severe limb and vocal paral- ysis,”Nature communications, vol. 13, no. 1, p. 6510, 2022

Show all 38 references
  1. [9]

    A high-performance speech neuroprosthesis,

    F. R. Willett, E. M. Kunz, C. Fan, D. T. Avansino, G. H. Wilson, E. Y . Choi, F. Kamdar, M. F. Glasser, L. R. Hochberg, S. Druck- mannet al., “A high-performance speech neuroprosthesis,”Na- ture, vol. 620, no. 7976, pp. 1031–1036, 2023

  2. [10]

    A bilingual speech neuroprosthesis driven by cortical articulatory representations shared between lan- guages,

    A. B. Silva, J. R. Liu, S. L. Metzger, I. Bhaya-Grossman, M. E. Dougherty, M. P. Seaton, K. T. Littlejohn, A. Tu-Chan, K. Gan- guly, D. A. Moseset al., “A bilingual speech neuroprosthesis driven by cortical articulatory representations shared between lan- guages,”Nature Biomed...

  3. [11]

    An accurate and rapidly calibrating speech neuroprosthe- sis,

    N. S. Card, M. Wairagkar, C. Iacobacci, X. Hou, T. Singer-Clark, F. R. Willett, E. M. Kunz, C. Fan, M. Vahdati Nia, D. R. Deo et al., “An accurate and rapidly calibrating speech neuroprosthe- sis,”New England Journal of Medicine, vol. 391, no. 7, pp. 609– 618, 2024

  4. [12]

    Speech syn- thesis from neural decoding of spoken sentences,

    G. K. Anumanchipalli, J. Chartier, and E. F. Chang, “Speech syn- thesis from neural decoding of spoken sentences,”Nature, vol. 568, no. 7753, pp. 493–498, 2019

  5. [13]

    Speech synthesis from ecog using densely connected 3d convolutional neural networks,

    M. Angrick, C. Herff, E. Mugler, M. C. Tate, M. W. Slutzky, D. J. Krusienski, and T. Schultz, “Speech synthesis from ecog using densely connected 3d convolutional neural networks,”Journal of neural engineering, vol. 16, no. 3, p. 036019, 2019

  6. [14]

    Gen- erating natural, intelligible speech from brain activity in motor, premotor, and inferior frontal cortices,

    C. Herff, L. Diener, M. Angrick, E. Mugler, M. C. Tate, M. A. Goldrick, D. J. Krusienski, M. W. Slutzky, and T. Schultz, “Gen- erating natural, intelligible speech from brain activity in motor, premotor, and inferior frontal cortices,”Frontiers in neuroscience, vol. 13, p. 1267, 2019

  7. [15]

    Real-time synthesis of imagined speech processes from mini- mally invasive recordings of neural activity,

    M. Angrick, M. C. Ottenhoff, L. Diener, D. Ivucic, G. Ivucic, S. Goulis, J. Saal, A. J. Colon, L. Wagner, D. J. Krusienskiet al., “Real-time synthesis of imagined speech processes from mini- mally invasive recordings of neural activity,”Communications bi- ology, vol. 4, no. 1,...

  8. [16]

    Online speech synthesis using a chronically implanted brain–computer interface in an individual with ALS,

    M. Angrick, S. Luo, Q. Rabbani, D. N. Candrea, S. Shah, G. W. Milsap, W. S. Anderson, C. R. Gordon, K. R. Rosenblatt, L. Claw- sonet al., “Online speech synthesis using a chronically implanted brain–computer interface in an individual with ALS,”Scientific reports, vol. 14, no....

  9. [17]

    Neuroincept decoder for high- fidelity speech reconstruction from neural activity,

    O. M. Khanday, J. L. P ´erez-C´ordoba, M. Y . Mir, A. A. Na- jar, and J. A. Gonzalez-Lopez, “Neuroincept decoder for high- fidelity speech reconstruction from neural activity,”arXiv preprint arXiv:2501.03757, 2025

  10. [18]

    A high-performance neuroprosthesis for speech decoding and avatar control,

    S. L. Metzger, K. T. Littlejohn, A. B. Silva, D. A. Moses, M. P. Seaton, R. Wang, M. E. Dougherty, J. R. Liu, P. Wu, M. A. Berger et al., “A high-performance neuroprosthesis for speech decoding and avatar control,”Nature, vol. 620, no. 7976, pp. 1037–1046, 2023

  11. [19]

    Silent speech inter- faces for speech restoration: A review,

    J. A. Gonzalez-Lopez, A. Gomez-Alanis, J. M. Mart ´ın Do ˜nas, J. L. P ´erez-C´ordoba, and A. M. Gomez, “Silent speech inter- faces for speech restoration: A review,”IEEE Access, vol. 8, pp. 177 995–178 021, 2020

  12. [20]

    Stable decoding from a speech bci enables control for an individual with ALS without recalibration for 3 months,

    S. Luo, M. Angrick, C. Coogan, D. N. Candrea, K. Wyse-Sookoo, S. Shah, Q. Rabbani, G. W. Milsap, A. R. Weiss, W. S. Ander- sonet al., “Stable decoding from a speech bci enables control for an individual with ALS without recalibration for 3 months,”Ad- vanced Science, vol. 10, ...

  13. [21]

    Longevity of a brain– computer interface for amyotrophic lateral sclerosis,

    M. J. Vansteensel, S. Leinders, M. P. Branco, N. E. Crone, T. Denison, Z. V . Freudenburg, S. H. Geukes, P. H. Gosselaar, M. Raemaekers, A. Schipperset al., “Longevity of a brain– computer interface for amyotrophic lateral sclerosis,”New Eng- land Journal of Medicine, vol. 391...

  14. [22]

    Relating EEG to continuous speech using deep neural networks: a review,

    C. Puffay, B. Accou, L. Bollens, M. J. Monesi, J. Vanthornhout, H. Van hamme, and T. Francart, “Relating EEG to continuous speech using deep neural networks: a review,”Journal of Neural Engineering, vol. 20, no. 4, p. 041003, 2023

  15. [23]

    The speech neuroprosthesis,

    A. B. Silva, K. T. Littlejohn, J. R. Liu, D. A. Moses, and E. F. Chang, “The speech neuroprosthesis,”Nature Reviews Neuro- science, vol. 25, no. 7, pp. 473–492, 2024

  16. [24]

    Brain–computer interfaces for restoring communi- cation,

    E. F. Chang, “Brain–computer interfaces for restoring communi- cation,”New England Journal of Medicine, vol. 391, no. 7, pp. 654–657, 2024

  17. [25]

    A neural speech decoding framework leveraging deep learning and speech synthesis,

    X. Chen, R. Wang, A. Khalilian-Gourtani, L. Yu, P. Dugan, D. Friedman, W. Doyle, O. Devinsky, Y . Wang, and A. Flinker, “A neural speech decoding framework leveraging deep learning and speech synthesis,”Nature Machine Intelligence, pp. 1–14, 2024

  18. [26]

    The potential of stereotactic-eeg for brain-computer interfaces: current progress and future directions,

    C. Herff, D. J. Krusienski, and P. Kubben, “The potential of stereotactic-eeg for brain-computer interfaces: current progress and future directions,”Frontiers in neuroscience, vol. 14, p. 123, 2020

  19. [27]

    Synthesiz- ing speech from intracranial depth electrodes using an encoder- decoder framework,

    J. Kohler, M. C. Ottenhoff, S. Goulis, M. Angrick, A. J. Colon, L. Wagner, S. Tousseyn, P. L. Kubben, and C. Herff, “Synthesiz- ing speech from intracranial depth electrodes using an encoder- decoder framework,”Neurons, Behavior, Data analysis, and The- ory, vol. 6, no. 1, dec 9 2022

  20. [28]

    Towards closed-loop speech synthesis from stereotactic EEG: A unit selection approach,

    M. Angrick, M. Ottenhoff, L. Diener, D. Ivucic, G. Ivucic, S. Goulis, A. J. Colon, L. Wagner, D. J. Krusienski, P. L. Kubben et al., “Towards closed-loop speech synthesis from stereotactic EEG: A unit selection approach,” inProc. ICASSP, 2022, pp. 1296–1300

  21. [29]

    Wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “Wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,”Advances in neural information processing systems, vol. 33, pp. 12 449–12 460, 2020

  22. [30]

    Language models are unsupervised multitask learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019

  23. [31]

    Toward a realistic model of speech processing in the brain with self-supervised learning,

    J. Millet, C. Caucheteux, Y . Boubenec, A. Gramfort, E. Dunbar, C. Pallier, J.-R. Kinget al., “Toward a realistic model of speech processing in the brain with self-supervised learning,”Advances in Neural Information Processing Systems, vol. 35, pp. 33 428– 33 443, 2022

  24. [32]

    Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,

    G. Tuckute, J. Feather, D. Boebinger, and J. H. McDermott, “Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,”Plos Biology, vol. 21, no. 12, p. e3002366, 2023

  25. [33]

    Do self- supervised speech and language models extract similar represen- tations as human brain?

    P. Chen, L. He, L. Fu, L. Fan, E. F. Chang, and Y . Li, “Do self- supervised speech and language models extract similar represen- tations as human brain?” inProc. ICASSP, 2024, pp. 2225–2229

  26. [34]

    Elastic net regulariza- tion paths for all generalized linear models,

    J. K. Tay, B. Narasimhan, and T. Hastie, “Elastic net regulariza- tion paths for all generalized linear models,”Journal of statistical software, vol. 106, 2023

  27. [35]

    Fasttext.zip: Compressing text classification models,

    A. Joulin, E. Grave, P. Bojanowski, M. Douze, H. J ´egou, and T. Mikolov, “Fasttext.zip: Compressing text classification models,” 2016. [Online]. Available: https://arxiv.org/abs/1612. 03651

  28. [36]

    Dataset of speech production in intracranial electroencephalography,

    M. Verwoert, M. C. Ottenhoff, S. Goulis, A. J. Colon, L. Wagner, S. Tousseyn, J. P. Van Dijk, P. L. Kubben, and C. Herff, “Dataset of speech production in intracranial electroencephalography,”Sci- entific data, vol. 9, no. 1, p. 434, 2022

  29. [37]

    The IFA corpus: A phonemically segmented dutch ‘open source’ speech database,

    R. van Son, D. Binnenpoorte, and L. Pols, “The IFA corpus: A phonemically segmented dutch ‘open source’ speech database,” inProceedings of Eurospeech. Data Archiving and Networked Services (DANS), Sep 2001

  30. [38]

    XLS-R: Self- supervised cross-lingual speech representation learning at scale,

    A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. V on Platen, Y . Saraf, J. Pinoet al., “XLS-R: Self- supervised cross-lingual speech representation learning at scale,” arXiv preprint arXiv:2111.09296, 2021

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.