Pith. sign in

REVIEW 1 cited by

LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of end-to-end ASR Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.02943 v1 pith:UL25MIH4 submitted 2023-10-04 cs.CL

classification cs.CL
keywords punctuationmodelscapitalizationbenchmarkend-to-endevaluationlibrispeech-pccapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traditional automatic speech recognition (ASR) models output lower-cased words without punctuation marks, which reduces readability and necessitates a subsequent text processing model to convert ASR transcripts into a proper format. Simultaneously, the development of end-to-end ASR models capable of predicting punctuation and capitalization presents several challenges, primarily due to limited data availability and shortcomings in the existing evaluation methods, such as inadequate assessment of punctuation prediction. In this paper, we introduce a LibriSpeech-PC benchmark designed to assess the punctuation and capitalization prediction capabilities of end-to-end ASR models. The benchmark includes a LibriSpeech-PC dataset with restored punctuation and capitalization, a novel evaluation metric called Punctuation Error Rate (PER) that focuses on punctuation marks, and initial baseline models. All code, data, and models are publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Universal-2-TF: Robust All-Neural Text Formatting for ASR

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Universal-2-TF performs punctuation restoration, truecasing, and inverse text normalization with a shared-encoder classifier followed by a span-level seq2seq model, beating the company's prior WFST-based system.

Pith tools