Pith. sign in

REVIEW 6 cited by

Retrieval Augmented Correction of Named Entity Speech Recognition Errors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.06062 v1 pith:NXZV36C3 submitted 2024-09-09 eess.AS cs.SD

classification eess.AScs.SD
keywords databaseentitiesentityerrorsqueriesrecognitionspeechsystems
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems, large language models (LLMs) have proven to be a versatile tool for various natural language processing (NLP) tasks. In NLP tasks where a database of relevant knowledge is available, retrieval augmented generation (RAG) has achieved impressive results when used with LLMs. In this work, we propose a RAG-like technique for correcting speech recognition entity name errors. Our approach uses a vector database to index a set of relevant entities. At runtime, database queries are generated from possibly errorful textual ASR hypotheses, and the entities retrieved using these queries are fed, along with the ASR hypotheses, to an LLM which has been adapted to correct ASR errors. Overall, our best system achieves 33%-39% relative word error rate reductions on synthetic test sets focused on voice assistant queries of rare music entities without regressing on the STOP test set, a publicly available voice assistant test set covering many domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DeRAGEC explicitly denoises retrieved named-entity candidates with phonetic scores, definitions, and synthetic rationales, improving ASR error-correction WER and NE hit ratio without additional training.

  2. Auto Review: Second Stage Error Detection for Highly Accurate Information Extraction from Phone Conversations

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An auto-review pipeline using multiple ASR transcript alternatives and LLM-generated pseudo-labels improves field extraction accuracy on healthcare benefit calls, but gains are inconsistent outside a fine-tuned model.

  3. LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context

    cs.SD 2025-05 conditional novelty 6.0 of 10

    An LLM-based ASR error corrector trained on synthetic rare-word speech and given simplified phonetic context lowers WER/CER and raises rare-word recall on English and Japanese benchmarks.

  4. A Theoretical Framework for Acoustic Neighbor Embeddings

    eess.AS 2024-12 conditional novelty 6.0 of 10

    Euclidean distances between acoustic neighbor embeddings are interpreted as phonetic similarity through a Bayes-error and Gaussian-isotropy approximation, with four validation experiments.

  5. GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems

    eess.AS 2025-01 conditional novelty 4.0 of 10

    GEC-RAG retrieves similar ASR/ground-truth examples via TF-IDF and uses them as few-shot prompts for GPT-4o, reporting large WER reductions for Persian.

  6. Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Fine-tuning Whisper on Estonian subtitles with iterative pseudo-labeling and test-time LLM editing improves subtitle quality, while LLM editing during training yields no gain.

Pith tools