REVIEW 7 cited by
Speak & Improve Challenge 2025: Tasks and Baseline Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents the "Speak & Improve Challenge 2025: Spoken Language Assessment and Feedback" -- a challenge associated with the ISCA SLaTE 2025 Workshop. The goal of the challenge is to advance research on spoken language assessment and feedback, with tasks associated with both the underlying technology and language learning feedback. Linked with the challenge, the Speak & Improve (S&I) Corpus 2025 is being pre-released, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The corpus consists of approximately 315 hours of audio data from second language English learners with holistic scores, and a 55-hour subset with manual transcriptions and error labels. The Challenge has four shared tasks: Automatic Speech Recognition (ASR), Spoken Language Assessment (SLA), Spoken Grammatical Error Correction (SGEC), and Spoken Grammatical Error Correction Feedback (SGECF). Each of these tasks has a closed track where a predetermined set of models and data sources are allowed to be used, and an open track where any public resource may be used. Challenge participants may do one or more of the tasks. This paper describes the challenge, the S&I Corpus 2025, and the baseline systems released for the Challenge.
Forward citations
Cited by 7 Pith papers
-
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
In two Transformer-based L2 speaking graders, a concept's linear recoverability in a hidden layer does not predict its influence on the predicted score, and sparse-autoencoder probing attenuates measured sensitivity.
-
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
Zero-shot LLM grading of L2 speech transcripts with CEFR descriptors beats a fine-tuned BERT baseline and matches a read-aloud-trained speech model on the S&I Corpus.
-
End-to-End Spoken Grammatical Error Correction
End-to-end Whisper models, trained with 2,500 hours of pseudo-labeled speech, fluent prompts, aligned references, and confidence filtering, outperform cascaded systems on spoken grammatical error correction and feedback.
-
Assessment of L2 Oral Proficiency using Speech Large Language Models
A speech LLM fine-tuned with a fair-average loss outperforms BERT and wav2vec2 baselines on holistic L2 oral proficiency scoring, and transfers across test parts and datasets.
-
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
Pseudo-labelling and prompting with fluent transcriptions improve end-to-end spoken grammatical error correction and feedback for Whisper-based models, but the benefits depend on model size.
-
Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems
Whisper trained with realistic 'um'/'uh' labels generated by Gemini reached 5.5% WER on L2 English speech, an 11.3% relative improvement over training with hesitations removed.
-
The NTNU System at the S&I Challenge 2025 SLA Open Track
Fusing a wav2vec 2.0 acoustic grader with a task-specific Phi-4 multimodal language model reduces RMSE to 0.375 on the L2 English speaking assessment challenge, ranking second.
Discussion (0). Continue with ORCID to comment.