Pith. sign in

REVIEW 3 cited by

STAB: Speech Tokenizer Assessment Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.02384 v1 pith:SRW6G7JE submitted 2024-09-04 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechtokenizerbenchmarkdownstreamframeworkstabtaskstokenizers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently, while several speech tokenizers have been proposed, there is ambiguity regarding the properties that are desired from a tokenizer for specific downstream tasks and its overall generalizability. Evaluating the performance of tokenizers across different downstream tasks is a computationally intensive effort that poses challenges for scalability. To circumvent this requirement, we present STAB (Speech Tokenizer Assessment Benchmark), a systematic evaluation framework designed to assess speech tokenizers comprehensively and shed light on their inherent characteristics. This framework provides a deeper understanding of the underlying mechanisms of speech tokenization, thereby offering a valuable resource for expediting the advancement of future tokenizer models and enabling comparative analysis using a standardized benchmark. We evaluate the STAB metrics and correlate this with downstream task performance across a range of speech tasks and tokenizer choices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking Prosody Encoding in Discrete Speech Tokens

    cs.SD 2025-08 unverdicted novelty 6.0 of 10

    A benchmark study shows which speech discretization choices preserve prosodic information, measured by how much token sequences change when prosody is artificially altered.

  2. SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

    eess.AS 2025-07 conditional novelty 5.0 of 10

    SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.

  3. Discrete Audio Representations for Automated Audio Captioning

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A supervised audio tokenizer trained with audio tagging labels nearly matches continuous BEATs features for automated audio captioning on Clotho, while conventional discrete tokens degrade performance.

Pith tools