REVIEW 6 cited by
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems, large language models (LLMs) have proven to be a versatile tool for various natural language processing (NLP) tasks. In NLP tasks where a database of relevant knowledge is available, retrieval augmented generation (RAG) has achieved impressive results when used with LLMs. In this work, we propose a RAG-like technique for correcting speech recognition entity name errors. Our approach uses a vector database to index a set of relevant entities. At runtime, database queries are generated from possibly errorful textual ASR hypotheses, and the entities retrieved using these queries are fed, along with the ASR hypotheses, to an LLM which has been adapted to correct ASR errors. Overall, our best system achieves 33%-39% relative word error rate reductions on synthetic test sets focused on voice assistant queries of rare music entities without regressing on the STOP test set, a publicly available voice assistant test set covering many domains.
Forward citations
Cited by 6 Pith papers
-
DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction
DeRAGEC explicitly denoises retrieved named-entity candidates with phonetic scores, definitions, and synthetic rationales, improving ASR error-correction WER and NE hit ratio without additional training.
-
Auto Review: Second Stage Error Detection for Highly Accurate Information Extraction from Phone Conversations
An auto-review pipeline using multiple ASR transcript alternatives and LLM-generated pseudo-labels improves field extraction accuracy on healthcare benefit calls, but gains are inconsistent outside a fine-tuned model.
-
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
An LLM-based ASR error corrector trained on synthetic rare-word speech and given simplified phonetic context lowers WER/CER and raises rare-word recall on English and Japanese benchmarks.
-
A Theoretical Framework for Acoustic Neighbor Embeddings
Euclidean distances between acoustic neighbor embeddings are interpreted as phonetic similarity through a Bayes-error and Gaussian-isotropy approximation, with four validation experiments.
-
GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems
GEC-RAG retrieves similar ASR/ground-truth examples via TF-IDF and uses them as few-shot prompts for GPT-4o, reporting large WER reductions for Persian.
-
Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs
Fine-tuning Whisper on Estonian subtitles with iterative pseudo-labeling and test-time LLM editing improves subtitle quality, while LLM editing during training yields no gain.
Discussion (0). Continue with ORCID to comment.