REVIEW 3 cited by
SEANet: A Multi-modal Speech Enhancement Network
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We explore the possibility of leveraging accelerometer data to perform speech enhancement in very noisy conditions. Although it is possible to only partially reconstruct user's speech from the accelerometer, the latter provides a strong conditioning signal that is not influenced from noise sources in the environment. Based on this observation, we feed a multi-modal input to SEANet (Sound EnhAncement Network), a wave-to-wave fully convolutional model, which adopts a combination of feature losses and adversarial losses to reconstruct an enhanced version of user's speech. We trained our model with data collected by sensors mounted on an earbud and synthetically corrupted by adding different kinds of noise sources to the audio signal. Our experimental results demonstrate that it is possible to achieve very high quality results, even in the case of interfering speech at the same level of loudness. A sample of the output produced by our model is available at https://google-research.github.io/seanet/multimodal/speech.
Forward citations
Cited by 3 Pith papers
-
Probing the Robustness Properties of Neural Speech Codecs
DAC is the most noise-robust neural codec at high bitrates, but at 3 kbps EnCodec wins, and measured non-linearity correlates with robustness.
-
CAPS: A Cascaded Reconstruction Model to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension
CAPS reconstructs wideband clean audio from 4 kHz, 8-bit sub-Nyquist hearable signals using a cascaded BWE plus multimodal SE network, cutting ADC power 3.3x while supporting 55 ms mobile inference.
-
Towards a Japanese Full-duplex Spoken Dialogue System
J-Moshi, the first public Japanese full-duplex spoken dialogue model, is built from Moshi and outperforms a Japanese dGSLM baseline on naturalness and meaningfulness.
Discussion (0). Continue with ORCID to comment.