Pith. sign in

REVIEW 2 cited by

The NCTE Transcripts: A Dataset of Elementary Math Classroom Transcripts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11772 v2 pith:WBPNAK5F submitted 2022-11-21 cs.CL

classification cs.CL
keywords classroomdatasettranscriptsinstructiondiscoursemovesscoresannotations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Classroom discourse is a core medium of instruction - analyzing it can provide a window into teaching and learning as well as driving the development of new tools for improving instruction. We introduce the largest dataset of mathematics classroom transcripts available to researchers, and demonstrate how this data can help improve instruction. The dataset consists of 1,660 45-60 minute long 4th and 5th grade elementary mathematics observations collected by the National Center for Teacher Effectiveness (NCTE) between 2010-2013. The anonymized transcripts represent data from 317 teachers across 4 school districts that serve largely historically marginalized students. The transcripts come with rich metadata, including turn-level annotations for dialogic discourse moves, classroom observation scores, demographic information, survey responses and student test scores. We demonstrate that our natural language processing model, trained on our turn-level annotations, can learn to identify dialogic discourse moves and these moves are correlated with better classroom observation scores and learning outcomes. This dataset opens up several possibilities for researchers, educators and policymakers to learn about and improve K-12 instruction. The dataset can be found at https://github.com/ddemszky/classroom-transcript-analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research

    cs.SD 2025-06 conditional novelty 6.0 of 10

    SimClass is a new 391-hour simulated classroom speech dataset with game-engine babble noise; ASR fine-tuning on it beats Librispeech and TEDLIUM on real classroom test sets.

  2. FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom Environments

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Fine-tuning pretrained speaker verification models on augmented children's speech reduces error rates in English-speaking classrooms, with ECAPA-TDNN nearly halving error on the MPT dataset.

Pith tools