Pith. sign in

REVIEW 3 cited by

Vec-Tok Speech: speech vectorization and tokenization for neural speech generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07246 v2 pith:KTRX56IM submitted 2023-10-11 cs.SD eess.AS

classification cs.SDeess.AS
keywords speechvec-tokgenerationhigh-fidelitylanguagemodelscodecgenerating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-fidelity texts or images in various tasks. In contrast, the current speech generative models are still struggling regarding speech quality and task generalization. This paper presents Vec-Tok Speech, an extensible framework that resembles multiple speech generation tasks, generating expressive and high-fidelity speech. Specifically, we propose a novel speech codec based on speech vectors and semantic tokens. Speech vectors contain acoustic details contributing to high-fidelity speech reconstruction, while semantic tokens focus on the linguistic content of speech, facilitating language modeling. Based on the proposed speech codec, Vec-Tok Speech leverages an LM to undertake the core of speech generation. Moreover, Byte-Pair Encoding (BPE) is introduced to reduce the token length and bit rate for lower exposure bias and longer context coverage, improving the performance of LMs. Vec-Tok Speech can be used for intra- and cross-lingual zero-shot voice conversion (VC), zero-shot speaking style transfer text-to-speech (TTS), speech-to-speech translation (S2ST), speech denoising, and speaker de-identification and anonymization. Experiments show that Vec-Tok Speech, built on 50k hours of speech, performs better than other SOTA models. Code will be available at https://github.com/BakerBunker/VecTok .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment

    eess.AS 2025-07 conditional novelty 6.0 of 10

    SemAlignVC strips source-speaker timbre by aligning a speech semantic encoder to BERT text embeddings, then resynthesizes the content conditioned only on a target voice reference.

  2. Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A linear residual that subtracts a speaker-embedding-based projection from WavLM representations reduces speaker information while preserving content, improving voice conversion.

  3. EchoFree: Towards Ultra Lightweight and Efficient Neural Acoustic Echo Cancellation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    EchoFree, a 278K-parameter hybrid echo canceller using Bark-scale features and a two-stage WavLM-guided training schedule, matches DeepVQE-S quality on the ICASSP 2023 AEC blind test.

Pith tools