Pith. sign in

REVIEW 10 cited by

Can Generative Large Language Models Perform ASR Error Correction?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04172 v2 pith:6KSY5GLI submitted 2023-07-09 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords correctionerrorgenerativelanguagemodelssystemapproachfashion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

ASR error correction is an interesting option for post processing speech recognition system outputs. These error correction models are usually trained in a supervised fashion using the decoding results of a target ASR system. This approach can be computationally intensive and the model is tuned to a specific ASR system. Recently generative large language models (LLMs) have been applied to a wide range of natural language processing tasks, as they can operate in a zero-shot or few shot fashion. In this paper we investigate using ChatGPT, a generative LLM, for ASR error correction. Based on the ASR N-best output, we propose both unconstrained and constrained, where a member of the N-best list is selected, approaches. Additionally, zero and 1-shot settings are evaluated. Experiments show that this generative LLM approach can yield performance gains for two different state-of-the-art ASR architectures, transducer and attention-encoder-decoder based, and multiple test sets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition

    eess.AS 2025-05 conditional novelty 7.0 of 10

    LibriSpeech and Common Voice evaluation sentences leak into the Pile, and controlled LLM pretraining experiments show that contamination biases output probabilities even when error rates barely change.

  2. DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    DeRAGEC explicitly denoises retrieved named-entity candidates with phonetic scores, definitions, and synthetic rationales, improving ASR error-correction WER and NE hit ratio without additional training.

  3. Customizing Speech Recognition Model with Large Language Model Feedback

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLM log-probability scores combined with acoustic scores serve as RL rewards to adapt ASR models to new domains without labeled data.

  4. LLM-based phoneme-to-grapheme for phoneme-based speech recognition

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Two-step phoneme-based ASR with LLM-P2G decoding, using noisy-phoneme augmentation and randomized top-K marginalized training, reduces WER on Polish and German versus WFST decoding.

  5. LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context

    cs.SD 2025-05 conditional novelty 6.0 of 10

    An LLM-based ASR error corrector trained on synthetic rare-word speech and given simplified phonetic context lowers WER/CER and raises rare-word recall on English and Japanese benchmarks.

  6. An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications

    eess.AS 2025-07 reject novelty 5.0 of 10

    The paper introduces AER, an LLM-judged question-answering metric for evaluating ASR output in LLM applications, and shows it does not correlate strongly with WER.

  7. PHRASED: Phrase Dictionary Biasing for Speech Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Phrase dictionary biasing, which matches source phrases in intermediate ASR text and then boosts or prompts the matching target phrases, improves phrase recall in streaming and LLM-based speech translation.

  8. Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A Data2Vec2 speech encoder pre-trained on 300,000 hours of unlabeled Chinese dialect speech, connected to a small Qwen LLM via a linear projector and fine-tuned in four stages, sets a new state of the art on Chinese d...

  9. Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

    cs.SD 2025-09 reject novelty 4.0 of 10

    An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.

  10. Large Language Models based ASR Error Correction for Child Conversations

    cs.CL 2025-05 conditional novelty 4.0 of 10

    LLM-based error correction improves zero-shot Whisper and fine-tuned WavLM child ASR transcriptions, but not fine-tuned Whisper, and conversational context as implemented degrades corrections.

Pith tools