REVIEW 5 cited by
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Benefiting from effective speech modeling, current Speech Large Language Models (SLLMs) have demonstrated exceptional capabilities in in-context speech generation and efficient generalization to unseen speakers. However, the prevailing information modeling process is encumbered by certain redundancies, leading to inefficiencies in speech generation. We propose Chain-of-Information Generation (CoIG), a method for decoupling semantic and perceptual information in large-scale speech generation. Building on this, we develop SpeechGPT-Gen, an 8-billion-parameter SLLM efficient in semantic and perceptual information modeling. It comprises an autoregressive model based on LLM for semantic information modeling and a non-autoregressive model employing flow matching for perceptual information modeling. Additionally, we introduce the novel approach of infusing semantic information into the prior distribution to enhance the efficiency of flow matching. Extensive experimental results demonstrate that SpeechGPT-Gen markedly excels in zero-shot text-to-speech, zero-shot voice conversion, and speech-to-speech dialogue, underscoring CoIG's remarkable proficiency in capturing and modeling speech's semantic and perceptual dimensions. Code and models are available at https://github.com/0nutation/SpeechGPT.
Forward citations
Cited by 5 Pith papers
-
Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models
Spoken math models that emit a 40%-compressed reasoning trace between question and answer beat full-reasoning baselines by ~3 accuracy points while using roughly one third of the text tokens.
-
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
XY-Tokenizer is a 1 kbps dual-channel speech codec that reports simultaneously strong text alignment and high speaker similarity, comparable to specialized codecs at similar bitrates.
-
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models
Ex-Omni is an OLLM that natively generates speech and ARKit-52 3D facial animation in one pass by using discrete speech units as temporal scaffolding and gated semantic injection.
-
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
Staged ASR-to-text-response-to-TTS training makes open end-to-end spoken dialogue systems trainable on 300 hours of public human-human data and more coherent than one-step speech-to-speech models.
-
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
OmniCharacter is a speech-language role-playing agent that generates character-specific voice responses with low latency, trained on a new 10K-dialogue, 135K-audio dataset of 20 game characters.
Discussion (0). Continue with ORCID to comment.