Pith. sign in

REVIEW 4 cited by

SpeechAlign: Aligning Speech Generation to Human Preferences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05600 v1 pith:SC5I3FJU submitted 2024-04-08 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modelslanguagespeechcodechumanspeechaligndistributionpreferences
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech language models have significantly advanced in generating realistic speech, with neural codec language models standing out. However, the integration of human feedback to align speech outputs to human preferences is often neglected. This paper addresses this gap by first analyzing the distribution gap in codec language models, highlighting how it leads to discrepancies between the training and inference phases, which negatively affects performance. Then we explore leveraging learning from human feedback to bridge the distribution gap. We introduce SpeechAlign, an iterative self-improvement strategy that aligns speech language models to human preferences. SpeechAlign involves constructing a preference codec dataset contrasting golden codec tokens against synthetic tokens, followed by preference optimization to improve the codec language model. This cycle of improvement is carried out iteratively to steadily convert weak models to strong ones. Through both subjective and objective evaluations, we show that SpeechAlign can bridge the distribution gap and facilitating continuous self-improvement of the speech language model. Moreover, SpeechAlign exhibits robust generalization capabilities and works for smaller models. Code and models will be available at https://github.com/0nutation/SpeechGPT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment

    cs.CL 2026-07 conditional novelty 6.0 of 10

    BoN TTS verifier rankings reverse across ASR families; same-family pairs recover 2–3× more oracle headroom, and cross-family rank ensembles give the most robust WER gains.

  2. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  3. Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Step-Audio-AQAA, a 130B end-to-end audio language model using dual-codebook tokens, text-audio interleaving, masked DPO and weight merging, is claimed to outperform Kimi-Audio and Qwen-Omni on the authors' StepEval-Au...

  4. A Preliminary Exploration with GPT-4o Voice Mode

    cs.CL 2025-02 conditional novelty 5.0 of 10

    GPT-4o voice mode is evaluated on 180 Dynamic-SUPERB tasks plus MMAU and CMM, showing strong audio understanding and low hallucination, but unstable refusal behavior and weak duration and instrument skills.

Pith tools