REVIEW 7 cited by
THCHS-30 : A Free Chinese Speech Corpus
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Speech data is crucially important for speech recognition research. There are quite some speech databases that can be purchased at prices that are reasonable for most research institutes. However, for young people who just start research activities or those who just gain initial interest in this direction, the cost for data is still an annoying barrier. We support the `free data' movement in speech recognition: research institutes (particularly supported by public funds) publish their data freely so that new researchers can obtain sufficient data to kick of their career. In this paper, we follow this trend and release a free Chinese speech database THCHS-30 that can be used to build a full- edged Chinese speech recognition system. We report the baseline system established with this database, including the performance under highly noisy conditions.
Forward citations
Cited by 7 Pith papers
-
PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
PARCO cuts named-entity errors by large margins (NE-CER 1.57% on AISHELL-1, NE-WER 8.34% on DATA2 at zero distractors) by combining phoneme-enriched entity encoding, a contrastive disambiguation loss, and hierarchical...
-
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
Mel-McNet performs online multichannel speech enhancement in the Mel domain, reducing FLOPs by roughly 60% versus McNet while keeping speech quality and ASR accuracy comparable.
-
Bridging the Data Provenance Gap Across Text, Speech and Video
A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...
-
Multilingual Speech Recognition with Corpus Relatedness Sampling
Corpus Relatedness Sampling, which anneals the training data distribution from uniform to target-focused based on cosine similarity of jointly learned corpus embeddings, outperforms fine-tuned multilingual baselines o...
-
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
The ASR-EC benchmark on Chinese ASR errors shows that multimodal LLM augmentation corrects ASR output best, while prompting alone worsens CER.
-
Speech-Driven End-to-End Language Discrimination towards Chinese Dialects
A speech-driven pipeline with MFCC features, HMM-DNN speech recognition, attention, and CNN fusion is presented for fine-grained Chinese dialect discrimination and evaluated on two benchmark corpora.
-
Inclusivity of AI Speech in Healthcare: A Decade Look Back
A decade-long audit finds persistent inclusivity gaps in speech AI for healthcare: English-heavy datasets, little demographic metadata, no speech-impaired samples, and limited bias research.
Discussion (0). Continue with ORCID to comment.