{"id":"56bb5c2a-115c-4520-b841-d7e5c806e55f","arxiv_id":"1908.00916","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A public test corpus of 39 noisy non-native English presentations with transcriptions and in-domain texts, on which three ASR baselines reach high word error rates.","lead":"The authors release a new speech test corpus of 39 short English presentations by European high school students, with transcriptions, slides, and web pages. The test set targets automatic speech recognition (ASR) on non-native, noisy speech, and the paper reports high word error rates for three baseline ASR systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth transcription quality is the load-bearing unvalidated premise: WERs in Table 7 are measured against a reference that was self-transcribed and only informally edited, with no independent verification.","rationale":"The reader identified the same weakest assumption: the accuracy of the ground-truth transcriptions in Section 2.3. My stress-test confirms that this is the single most load-bearing premise, because the paper's empirical contribution is the WER evaluation in Section 4, and WER is computed against exactly those transcriptions. The manuscript itself asserts that students had incentive and that the authors reviewed the transcripts, but it provides no independent verification, inter-annotator agreement, or alternate-reference analysis. This is a concrete, testable gap rather than a fatal flaw: high WERs of 45-90% would almost certainly remain qualitatively high even with a cleaner reference, but precise benchmark numbers and corpus reusability depend on reference quality. I therefore recommend a conditional acceptance: the resource is valuable, but the authors should either supply an independent transcription-quality check or explicitly characterize the WER figures as indicative pending such verification. This does not require rejection, and it does not change the assessment of novelty or the fundamental usefulness of the corpus.","tokens_in":7355,"tokens_out":2380,"duration_ms":28259,"concrete_test":"Select a random subset of at least 10 recordings from the released corpus and have two independent transcribers, blind to the original transcripts and to each other, produce verbatim reference transcriptions under a written protocol that includes fillers, false starts, and disfluencies. Then compute (a) the word error rate between the released reference and each independent transcription, treating the independent transcription as the reference, and (b) the WER of one baseline ASR system, e.g. JRTk as reported in Table 7, computed against both the released and the independent transcriptions. If the released-to-independent reference discrepancy exceeds roughly 5% absolute WER on average, the released reference and Table 7 numbers require revision or a documented caveat; if it is below that threshold, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim of the paper is that the corpus is a usable ASR test set, and the reported evidence for this is the baseline WER evaluation in Section 4. That evaluation depends entirely on the reference transcriptions: Section 4.2 states that the transcription obtained from participants was taken as ground truth. The quality-control procedure described in Section 2.3 is that students transcribed their own speech and the authors subsequently reviewed and edited the transcripts, but no independent transcription, no inter-annotator agreement, and no quantitative check of reference accuracy is reported. Participant self-transcription can systematically normalize the speech: speakers may omit fillers, false starts, and repeated words, or write what they intended to say rather than what was actually said. The authors explicitly preserved non-standard grammar and vocabulary, which is defensible for authenticity but further removes the reference from a strictly verbatim standard. If the reference is systematically non-verbatim or contains transcription errors, then the WER values in Table 7 are biased and the benchmark is unreliable for comparing ASR systems, even though the qualitative picture of high error rates would likely survive. The corpus value as a resource depends on this premise, and it is currently unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a new, publicly released test corpus composed of 39 audio recordings of practice business presentations given by high-school students in English as a second language, together with manual transcriptions, presentation slides, and web pages of the students' fictional firms. The authors describe the collection methodology, including an incentive scheme in which participants transcribed their own speech, and report baseline word error rates for three ASR systems (Google Cloud Speech-to-Text, Kaldi BBC, and JRTk). The reported WERs are high, e.g., a mean WER of 45.63% for JRTk on all recordings, and the paper positions the corpus as a challenging test set for ASR and domain adaptation.","tokens_in":7523,"tokens_out":7547,"duration_ms":72084,"significance":"If the resource is reliable, it fills a genuine gap: a publicly available, noisy L2-English speech test set with in-domain texts and named entities, collected under realistic trade-fair conditions with eight European L1 backgrounds. The authors are careful about ethical compliance and data formats, and the corpus is released with a persistent handle, which is commendable. The baseline WER measurements, however, are trustworthy only to the extent that the reference transcriptions are verbatim and accurate; the paper currently provides no independent verification of transcription quality. The resource itself has value independent of the baseline numbers, but the benchmark claims in Section 4 need strengthening before the numbers can be used as a reliable comparison point.","major_comments":[{"comment":"The WER benchmark in Section 4 is scored against reference transcripts that were produced by the speakers themselves and then informally edited by the authors (Section 2.3), with no independent transcription, inter-annotator agreement, or quantitative verification of verbatim accuracy. Because speakers may systematically normalize their own speech (e.g., omitting fillers, false starts, or repairs, or writing what they intended rather than what was actually uttered), the reference may be non-verbatim in ways that bias the reported WERs. Section 4.2 states that the participant transcripts were taken as ground truth without further validation. The paper should either report an independent verification (e.g., re-transcribe a sample of the recordings and provide agreement statistics) or explicitly qualify the reported WER values as preliminary and state this limitation as a caveat for users of the corpus.","section":"Section 2.3 and Section 4.2"},{"comment":"In Table 7, the Kaldi BBC model's WER has an implausibly narrow spread across the corpus (standard deviation 2.29 on the 'Recognized by all' subset, min/max 83.96/91.03 on all recordings), while Google and JRTk vary by tens of percentage points. This pattern suggests the Kaldi outputs may be dominated by a systematic artifact, such as a constant misrecognition or a decoding failure that still returns text. The authors should provide a few representative Kaldi transcripts alongside their references to rule out a degenerate decoding path, and they should report the number of recordings in the 'Recognized by all' subset so the reader can assess how much data the left-hand columns of Table 7 represent.","section":"Section 4.3, Table 7"}],"minor_comments":[{"comment":"The single-speaker row of Table 3 lists counts that sum to 15 rather than the stated total of 17; the per-language counts (cs, es, ro, sk, hu) should be checked and corrected.","section":"Table 3"},{"comment":"The Web row of Table 6 sums to 24 rather than the listed total of 23; this should be reconciled with Table 5, which implies 24 firms with web pages (20 with both slides and web plus 4 with web only).","section":"Table 6"},{"comment":"The abstract contains the non-standard word 'beneﬁtable', presumably a typo for 'beneficial', and Section 2.2 says winners 'were awarded prices for their performances', which should read 'prizes'.","section":"Abstract and Section 2.2"},{"comment":"The sentence 'We also tried Microsoft Cloud ASR but it failed for all our recordings' should specify the failure mode (empty outputs, server errors, or other) and the number of recordings attempted.","section":"Section 4.1"},{"comment":"The paper should report how many recordings fall in the 'Recognized by all' subset, since the left-hand columns of Table 7 and Figure 1 are based on an unreported number of data points.","section":"Section 4.3"},{"comment":"Reference [17] contains typographical errors in the author names ('Mller', 'Stker', 'Zenkel') and should be corrected to the proper spellings (Müller, Stüker, Zenkel).","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a useful corpus with a clear release, but the evaluation section currently rests on unvalidated reference transcriptions and on a Kaldi result whose uniformity I find suspicious. If the authors provide a small independent verification study or, at minimum, a clear limitation statement, and if they explain the Kaldi behavior, I would support acceptance; otherwise the benchmark claims should be treated with caution. The resource itself, independent of the baseline numbers, seems worth publishing after such clarifications."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a modest but genuinely useful corpus paper. The new thing is the package: L2 English from eight European L1s, realistic noisy fairground audio, and aligned slides and web pages for 36 of the 39 firms. No prior speech corpus I know of combines those three elements. The release itself is the contribution, and it is a reasonable one for ASR evaluation of non-native speech in noisy conditions.\n\nThe paper also does several things well. The collection methodology is clearly described, including the incentive structure for the transcription competition. The authors are transparent about the messy parts: one ASR system failed completely, Google returned empty outputs on some recordings, and the recordings contain live music and announcements. They preserved non-standard grammar in the transcripts instead of silently \"correcting\" it, which is the right choice for authenticity. The baseline WERs are presented as descriptive statistics, not as a scientific claim about ASR limitations. For a corpus paper, that is appropriate.\n\nThe soft spot is exactly where the stress-test note points. The ground-truth transcripts were produced by the student speakers themselves and then edited by the authors. There is no independent transcription, no inter-annotator agreement, and no quantitative check of reference accuracy. Section 4.2 simply takes the participant transcripts as ground truth. Self-transcription can systematically normalize speech: speakers omit fillers, false starts, and repeated words, or write what they intended to say rather than what they actually said. If that happened here, every WER in Table 7 is biased, and the benchmark is less reliable for comparing ASR systems. I think the qualitative picture would survive—this audio is clearly hard for current ASR—but the precise numbers should not be treated as a rigorous benchmark until the reference is verified.\n\nA second, minor point: the paper claims the additional texts could help domain adaptation, but does not demonstrate that they actually do. That is fine for a resource paper, as long as the claim stays a suggestion rather than a result.\n\nWho is this for? ASR researchers working on non-native English, noisy conditions, or domain adaptation. They will get value from the corpus itself and from the honest baseline documentation. The paper deserves a serious referee, but the authors should be asked to either add some independent verification of a sample of transcriptions or explicitly label the WERs as approximate and frame the corpus as a challenge set rather than a precisely labeled benchmark.","headline":"Worth a look: a small but real L2 English ASR test set with noisy audio and aligned slides/web texts, though its headline WERs rest on unverified self-transcriptions and should be treated as approximate.","tokens_in":8106,"tokens_out":1620,"would_cite":true,"duration_ms":17788,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents a publicly released test corpus of 39 noisy, non-native English business presentations with slides and webpages, and reports that three baseline ASR systems all fail on it, with the best system's mean word error rate…","keywords":["automatic speech recognition","ASR evaluation","speech corpus","non-native English","L2 English","business presentations","noisy speech","domain adaptation"],"falsifier":"Have an independent professional transcriber re-transcribe a sample of the recordings, and compute the word error rate between the corpus reference and the professional transcript; if that disagreement is as large as the ASR error rates (tens of percent), then the ground-truth assumption fails and the benchmark cannot separate ASR errors from transcript errors.","tokens_in":7162,"feed_emoji":"🎤","tokens_out":6962,"duration_ms":64648,"temperature":0.7,"pith_summary":"The paper's aim is to give the speech-recognition community a test set that matches a real, difficult setting: short English business pitches given by non-native speakers in a noisy hall. The corpus contains 39 presentations, 61 speakers, about one hour of audio, and reference transcriptions written by the speakers and then edited by the authors. For 36 of the firms, it also includes the presentation slides and web pages, so a system can be tested with and without access to in-domain vocabulary and named entities before hearing the talk. The authors benchmark three ASR systems and show that none handles the data well: the best mean word error rate is 45.63%.","feed_headline":"New speech test set gives best ASR 45% word error","feed_subtitle":"It pairs noisy student business talks with slides and webpages, testing whether in-domain text helps speech recognition.","key_machinery":"The load-bearing object is the test corpus itself, built as triplets: audio recordings, speaker-produced reference transcripts, and in-domain texts (slides, web pages) for 36 of 39 firms. The authors define the evaluation setup through word error rate (WER)—the minimum number of insertions, deletions, and substitutions needed to edit the ASR output into the reference, divided by reference word count—computed with case and punctuation ignored. The baselines are defined by their training resources (TED-LIUM 3 and Broadcast News for JRTk; 1600 hours of BBC audio and subtitle text for the Kaldi model), so the corpus's reported failure rates are tied to those configurations. The additional texts come in three formats (original, XLIFF, plaintext), which the authors suggest makes them easy to use for vocabulary extraction or adaptation experiments.","core_discovery":"The central claim is that a one-hour corpus of 39 recordings of student-run business presentations, spoken in L2 English by 61 European high-school students and captured with headset microphones in a noisy trade-fair environment, is a usable and challenging public test set for ASR. Its special feature is that extra relevant texts—slides and web pages of the fictional companies—are packaged in original, XLIFF, and plaintext forms, giving evaluators a way to study whether supplying in-domain vocabulary and named entities before recognition improves output. The paper establishes the difficulty by evaluating three baselines: a JRTk system trained on TED talks and Broadcast News, a Kaldi model trained on BBC broadcast data, and Google Cloud Speech-to-Text. On all 39 recordings, mean word error rates are 45.63% for JRTk, 89.32% for Google, and 87.47% for the Kaldi-BBC system, with JRTk's individual scores ranging from 25% to over 99%; the authors take this as evidence that current systems are far from robust on accented, noisy, spontaneous speech.","pith_inferences":["A natural extension not reported in the paper is a controlled adaptation experiment: take the slides and web pages, extract their named entities, bias the ASR language model, and measure the change in WER specifically on entity words; the corpus's design supports exactly this comparison.","Because the reference transcripts deliberately preserve non-standard learner grammar and vocabulary, the corpus measures ASR against authentic L2 speech, not idealized English; systems fine-tuned on corrected transcripts may look worse on it than they would in a deployment where speakers' errors are accepted.","The one-hour size makes the corpus unsuitable for training, which is likely the authors' intent; its value is as a targeted test set, and combining it with larger training corpora in a multi-task setting is a plausible use."],"forward_implications":["The corpus offers a ready-made benchmark for noisy, non-native, spontaneous English, with per-recording WER spread wide enough to distinguish robust systems from brittle ones.","Because slides and web pages contain the same named entities and domain words as the talks, the corpus allows a controlled test of whether exposing an ASR system to that text beforehand reduces errors.","The reported failure of a cloud ASR system on some recordings (100% WER from empty outputs) indicates that the corpus can also stress-test a system's noise robustness, not just its language model.","The corpus gives a way to evaluate whether ASR models trained on L1 English speech, such as TED talks or broadcast audio, transfer to the European L2 English found in international business settings."],"supporting_citations":[{"why":"Provides the JRTk speech recognition toolkit used as the best-performing baseline system.","marker":"[9]"},{"why":"Supplies the IBIS single-pass decoder used by the JRTk baseline to produce its outputs.","marker":"[15]"},{"why":"Gives the TED-LIUM 3 corpus on which the JRTk acoustic model was trained.","marker":"[6]"},{"why":"Gives the 1996 Broadcast News corpus, the second training source for the JRTk acoustic model.","marker":"[4]"},{"why":"Describes the IWSLT 2017 lecture-transcription system, the exact JRTk configuration evaluated here.","marker":"[17]"},{"why":"Provides the Kaldi toolkit underlying the Kaldi-BBC baseline.","marker":"[12]"},{"why":"Supplies the Multi-Genre Broadcast Challenge data (1600 hours of BBC audio, subtitle text) on which the Kaldi-BBC baseline was trained.","marker":"[2]"}],"fun_headline_variants":["Student pitch test: ASR systems hit 45-89% word error","Noisy student talks stump ASR in new benchmark set","Speech test set pairs slides with talks, WER hits 89%","ASR flubs accented student pitches: WER up to 89%","New corpus: 39 student presentations stress ASR systems"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported word error rates are only meaningful if the reference transcriptions, written by the student speakers and lightly edited by the authors, accurately represent what was said; if those transcripts are systematically incomplete or paraphrased, every WER in the paper would be off.","fun_headline_variants_meta":{"raw":{"variants":["Student pitch test: ASR systems hit 45-89% word error","Noisy student talks stump ASR in new benchmark set","Speech test set pairs slides with talks, WER hits 89%","ASR flubs accented student pitches: WER up to 89%","New corpus: 39 student presentations stress ASR systems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1369,"prompt_tokens":854,"completion_tokens":515,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":422}},"tokens_in":470,"tokens_out":515,"duration_ms":6309,"temperature":1.0,"reasoning_tokens":422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:27:12.627108+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent professional transcriber re-transcribe a sample of the recordings, and compute the word error rate between the corpus reference and the professional transcript; if that disagreement is as large as the ASR error rates (tens of percent), then the ground-truth assumption fails and the benchmark cannot separate ASR errors from transcript errors.","supporting_citations":[{"cited_title":"In: P roceedings of ICASSP 97 (Jan 1997)","cited_arxiv_id":null,"evidence_quote":"Provides the JRTk speech recognition toolkit used as the best-performing baseline system."},{"cited_title":"In: IEEE Works hop on Automatic Speech Recognition and Understanding, 2001","cited_arxiv_id":null,"evidence_quote":"Supplies the IBIS single-pass decoder used by the JRTk baseline to produce its outputs."},{"cited_title":"In: Proceedings of the 1997 DARP A Speech Recognition Worksh op","cited_arxiv_id":null,"evidence_quote":"Gives the 1996 Broadcast News corpus, the second training source for the JRTk acoustic model."},{"cited_title":"In: The International Workshop on Spoken Lan- guage Translation (IWSLT)","cited_arxiv_id":null,"evidence_quote":"Describes the IWSLT 2017 lecture-transcription system, the exact JRTk configuration evaluated here."},{"cited_title":"In: IEEE 2011 Work shop on Auto- matic Speech Recognition and Understanding","cited_arxiv_id":null,"evidence_quote":"Provides the Kaldi toolkit underlying the Kaldi-BBC baseline."},{"cited_title":"In: Proc","cited_arxiv_id":null,"evidence_quote":"Supplies the Multi-Genre Broadcast Challenge data (1600 hours of BBC audio, subtitle text) on which the Kaldi-BBC baseline was trained."}],"review_version":1}