Pith. sign in

REVIEW 2 cited by

Digital Voicing of Silent Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02960 v1 pith:RVLIBFMA submitted 2020-10-06 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords silentspeechvocalizedaudiocollecteddataduringmeasurements
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we consider the task of digitally voicing silent speech, where silently mouthed words are converted to audible speech based on electromyography (EMG) sensor measurements that capture muscle impulses. While prior work has focused on training speech synthesis models from EMG collected during vocalized speech, we are the first to train from EMG collected during silently articulated speech. We introduce a method of training on silent EMG by transferring audio targets from vocalized to silent signals. Our method greatly improves intelligibility of audio generated from silent EMG compared to a baseline that only trains with vocalized data, decreasing transcription word error rate from 64% to 4% in one data condition and 88% to 68% in another. To spur further development on this task, we share our new dataset of silent and vocalized facial EMG measurements.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Brain-to-Text Decoding with Context-Aware Neural Representations and Large Language Models

    eess.SP 2024-11 conditional novelty 6.0 of 10

    Diphone-based marginalization plus LLM refinement achieves state-of-the-art word error rate (5.77%) and phoneme error rate (15.34%) on the Brain-to-Text 2024 benchmark.

  2. From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A transformer ASR with GPT-2 post-correction reduces word error rate for EMG-based silent speech recognition from 36% to 30% on the Digital Voicing test set.

Pith tools