Pith. sign in

REVIEW 3 cited by

My Science Tutor (MyST) -- A Large Corpus of Children's Conversational Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.13347 v1 pith:IOHO7R47 submitted 2023-09-23 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords corpuschildrenconversationalmystsciencespeechtutorapproximately
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This article describes the MyST corpus developed as part of the My Science Tutor project -- one of the largest collections of children's conversational speech comprising approximately 400 hours, spanning some 230K utterances across about 10.5K virtual tutor sessions by around 1.3K third, fourth and fifth grade students. 100K of all utterances have been transcribed thus far. The corpus is freely available (https://myst.cemantix.org) for non-commercial use using a creative commons license. It is also available for commercial use (https://boulderlearning.com/resources/myst-corpus/). To date, ten organizations have licensed the corpus for commercial use, and approximately 40 university and other not-for-profit research groups have downloaded the corpus. It is our hope that the corpus can be used to improve automatic speech recognition algorithms, build and evaluate conversational AI agents for education, and together help accelerate development of multimodal applications to improve children's excitement and learning about science, and help them learn remotely.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SimClass is a new 391-hour simulated classroom speech dataset with game-engine babble noise; ASR fine-tuning on it beats Librispeech and TEDLIUM on real classroom test sets.

  2. Causal Analysis of ASR Errors for Children: Quantifying the Impact of Physiological, Cognitive, and Extrinsic Factors

    eess.AS 2025-02 reject novelty 6.0 of 10

    For children's ASR, word error rates are driven most by utterance length and child age, then noise and pronunciation, and fine-tuning lowers age sensitivity but not length sensitivity.

  3. Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Acoustic-focused augmentation of a 960-hour dataset is reported to reduce out-of-distribution word error rates by up to 19.24 percent, suggesting acoustic diversity, not linguistic diversity, drives ASR robustness.

Pith tools