Pith. sign in

REVIEW 2 cited by

N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.00456 v3 pith:SXODIEYQ submitted 2023-03-01 cs.CL cs.SDeess.AS

N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space

classification cs.CL cs.SDeess.AS
keywords correctionmodeln-bestdecodingerrorinputconstrainedinformation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Error correction models form an important part of Automatic Speech Recognition (ASR) post-processing to improve the readability and quality of transcriptions. Most prior works use the 1-best ASR hypothesis as input and therefore can only perform correction by leveraging the context within one sentence. In this work, we propose a novel N-best T5 model for this task, which is fine-tuned from a T5 model and utilizes ASR N-best lists as model input. By transferring knowledge from the pre-trained language model and obtaining richer information from the ASR decoding space, the proposed approach outperforms a strong Conformer-Transducer baseline. Another issue with standard error correction is that the generation process is not well-guided. To address this a constrained decoding process, either based on the N-best list or an ASR lattice, is used which allows additional information to be propagated.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Voice Memory for Agentic Speech Recognition

    cs.CL 2026-07 conditional novelty 6.0

    Score-gated text memories let a frozen LLM corrector decide when not to edit ASR hypotheses, cutting weighted WER from 8.36% to 7.52% without weight updates.

  2. Non-Intrusive Automatic Speech Recognition Refinement: A Survey

    eess.AS 2025-08 accept novelty 4.0

    A survey that classifies non-intrusive ASR refinement methods into five categories, reviews domain adaptation and evaluation datasets, proposes standardized metrics, and identifies future research directions.