{"id":"ece65e4b-f064-4ffe-9fbe-ec98d0baad79","arxiv_id":"2412.11538","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An openly released 630M-parameter BEST-RQ speech encoder, pretrained on 200k hours, matches state-of-the-art encoders on several ASR benchmarks and holds its own on ten SUPERB tasks.","lead":"MERaLiON-SpeechEncoder is a 630-million-parameter speech model pretrained on 200,000 hours of English-heavy audio using the BEST-RQ self-supervised objective. It is released openly and shows competitive speech recognition on spontaneous and Singapore-accented English, with a public evaluation across ten SUPERB tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"NSC test-set overlap with the 10k-hour pretraining corpus is unstated; if test audio appears unlabelled in pretraining, the headline Singapore ASR gains are inflated.","rationale":"The reader's weakest assumption is exactly the load-bearing point. The Singapore-English result is the paper's most distinctive contribution, and it is the only result that cannot be checked against a public benchmark without knowing the pretraining/test partition boundary. The SUPERB test-set hyperparameter selection is a real flaw, but it is secondary: it affects the 'competitive across ten tasks' claim without invalidating the core ASR result, and the qualitative conclusion is unlikely to flip. The TEDLIUMv3 'spontaneous speech improvement' is also weakened by the stronger zero-shot Whisper baseline, but the paper compares against a same-setup WavLM finetune, so that claim is internally consistent. The NSC leakage issue, by contrast, is a single unstated exclusion that, if violated, would directly invalidate the reported Singapore gains. My proposed check—manifest-level overlap comparison—would settle it. If the authors confirm exclusion and publish manifests, the condition is satisfied; if not, the Singapore claim should be downgraded or the experiments rerun. The reader's CONDITIONAL verdict already captures this, so I recommend no change.","tokens_in":15016,"tokens_out":6080,"duration_ms":55689,"concrete_test":"Obtain the NSC pretraining manifest (or the data-loading script) and the NSC evaluation splits from Appendix A.2/Table 9; match audio file paths or hashes across the six test parts. If any overlap is found, rerun the NSC ASR finetuning with pretraining restricted to the post-filtering train split and recompute Table 7; if the Part 2 or average WER gaps narrow materially, the Singapore ASR claim is unsupported. If manifests cannot be released, a minimal acceptable check is a written statement plus code proving test files were excluded at the pretraining data-loading stage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that MERaLiON improves Singapore-English ASR depends on the NSC test utterances (Table 9, 37.65 h total) being absent from the 10k h NSC component of the pretraining corpus (Table 1). The paper never states this. Appendix A.2's 'Data Splits Consistency' note only ensures identical transcriptions are in the same supervised split; it does not say the unsupervised pretraining manifest excluded the test set. There is also an unresolved quantity: Table 1 lists NSC as 10k h, while Table 9 gives 8169.48 h train + 37.65 h test after filtering, so it is unclear whether pretraining used the filtered train split or the full raw corpus. If any test audio appeared unlabelled in BEST-RQ pretraining, the encoder saw the exact waveforms and can encode utterance- or speaker-specific acoustic structure; CTC finetuning can then exploit that, inflating the reported Part 1-6 WERs (7.2/11.5/20.4/27.4/14.4/10.4, avg 15.2) and the 36.5% relative gain on Part 2. Since each Part test split is only 4-8 h, even small overlap can materially shift these numbers, so the Singapore contribution is conditional on an unverified exclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MERaLiON-SpeechEncoder, a 630M-parameter Conformer-based speech encoder pretrained with the BEST-RQ objective on roughly 200k hours of speech, initialized from a Libri-light checkpoint and then continuously pretrained on an expanded corpus that includes 10k hours of Singapore English from the National Speech Corpus (NSC). The authors report ASR results after CTC finetuning on LibriSpeech, TEDLIUMv3, and a 420-hour NSC subset, as well as results on ten SUPERB tasks. They claim improvements over an in-house finetuned WavLM large on spontaneous speech and Singapore English, and competitive performance with Wav2Vec 2.0, HuBERT, and WavLM across SUPERB. The model checkpoints are released on Hugging Face.","tokens_in":15267,"tokens_out":4424,"duration_ms":41953,"significance":"If the reported results are taken at face value, the paper's main contribution is an open, independently implemented, large-scale BEST-RQ encoder with detailed training and evaluation documentation. This fills a real gap, since no public BEST-RQ model at this scale has been released, and the focus on Singapore English/Singlish is valuable for the regional speech community. The detailed account of AMD-GPU pretraining and the SUPERB evaluation beyond ASR are also useful. However, the central Singapore-ASR claim and the SUPERB comparison rest on data-hygiene issues that must be resolved before the results can be interpreted as stated.","major_comments":[{"comment":"The Singapore-English claim depends on whether the NSC evaluation utterances were excluded from the 10k-hour NSC component of the pretraining corpus. The paper states in Table 1 that NSC supplied about 10k hours of unlabelled pretraining data, while Table 9 shows that after filtering there are 8169.48 hours of NSC train data and 37.65 hours of NSC test data. The text never states that the Table 9 test utterances were withheld from pretraining. Appendix A.2's 'Data Splits Consistency' note only prevents identical transcriptions from appearing in different supervised splits; it does not address the unsupervised pretraining manifest. Since BEST-RQ pretraining consumes raw waveforms without labels, any overlap between pretraining audio and the test set would let the CTC finetuning stage exploit utterance- or speaker-specific acoustic structure. Each NSC Part test split is only 4–8 hours, so even modest contamination could materially inflate the reported WERs and the 36.5% relative gain on Part 2. Please state explicitly whether the NSC test subset was excluded from the pretraining corpus, and if the full raw NSC corpus was used, provide the exact manifest or filtering procedure that guarantees exclusion.","section":"Section 3.3, Table 1, and Appendix A.2/Table 9"},{"comment":"The SUPERB comparison is weakened by test-set selection of hyperparameters. Section 5.2 states that the learning rates and batch sizes in Table 10 'were chosen simply by using the best result on the test set, between either the default SUPERB hyperparameters or those chosen for WavLM large.' Selecting hyperparameters on the test set inflates the reported SUPERB scores relative to baselines whose numbers may come from default settings. Please report which of the two configurations was chosen for each task, present results for both configurations, or select hyperparameters on the development set. This is load-bearing for the claim that the encoder is 'competitive' with other state-of-the-art encoders across ten SUPERB tasks.","section":"Section 5.2, Table 8, and Table 10"}],"minor_comments":[{"comment":"The abstract and introduction say the model was 'pre-trained from scratch on 200,000 hours,' but Section 3.4 explains that the full pretraining was initialized from a checkpoint pretrained on Libri-light (60k hours). Please rephrase to avoid the contradictory 'from scratch' description.","section":"Abstract, Section 1, and Section 3.4"},{"comment":"The row for 'Whisper large v3 (finetuned in-house)' appears as '81694.4' in the Finetuning column; this should be '8169' followed by the separate WER '4.4'.","section":"Table 7"},{"comment":"There are several typographical and spacing issues, including 'V oxPopuli' and 'V oice' for VoxPopuli and Voice, 'retraining' instead of 'retaining' in the contributions list, and 'non-conclusive' instead of 'non-exhaustive' in Section 6. These should be corrected.","section":"Throughout"},{"comment":"The random selection of 70 hours from each NSC part is not reproducible as described; please provide the random seed or a stable identifier for the chosen subset, including the test subset construction.","section":"Appendix A.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a technical report with an engineering contribution rather than a new learning method, but that can be appropriate for this venue if the empirical claims are reliable. The NSC pretraining/evaluation overlap is the deciding issue: without an explicit statement or manifest proving exclusion of the test set from pretraining, the headline Singapore-English result cannot be accepted. The SUPERB test-set hyperparameter selection is a secondary but still important fix. I would ask the authors to address both in a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a useful technical report with a genuinely valuable artifact—the first openly released BEST-RQ speech encoder at scale (630M, 200k hours) plus a SUPERB evaluation that gives the community its first look at how random-projection quantization transfers beyond ASR. The model is on Hugging Face, the ablations on masking probability, codebook count/vocabulary size, and positional embeddings are concrete, and the in-house WavLM large finetuning baseline makes the comparisons fair. The authors also clearly admit WavLM still wins overall on SUPERB, which is the right kind of honesty for a model release.\n\nThe main soft spot is the Singapore English claim. The NSC test set (37.65 h across six parts, Table 9) is not stated to be excluded from the 10 k h NSC component of the pretraining data (Table 1). The Appendix's 'Data Splits Consistency' note only covers supervised finetuning splits, not the unsupervised pretraining manifest. If any test audio appeared unlabelled during BEST-RQ pretraining, the encoder could memorize utterance- or speaker-specific acoustic patterns that CTC finetuning can exploit. Each part's test split is only 4–8 h, so even small overlap can shift the reported 36.5% relative gain on Part 2. This isn't proof of leakage, but the paper needs one explicit sentence ruling it out before the headline 'improvements to Singapore speech' is taken at face value.\n\nTwo smaller issues. First, the SUPERB hyperparameters were chosen using test-set performance (between two configurations), which mildly inflates those numbers; not a big deal, but worth disclosing more prominently. Second, the abstract says they improve spontaneous speech, yet the stronger Whisper large v2 zero-shot baseline scores 4.0 WER on TEDLIUMv3 versus their 5.6; the body mentions it, the abstract doesn't, which overstates the finding.\n\nThe math and citations are clean: no fitted-parameter circularity, the BEST-RQ implementation is independent, and the self-citation to MERaLiON-AudioLLM doesn't carry any evidential weight. The paper is a solid technical report, not a breakthrough.\n\nWho it's for: people building or benchmarking speech encoders, especially for Singapore English and code-switched speech; also anyone curious about BEST-RQ at scale. It deserves peer review, but the revision should clarify the NSC pretraining/test split and fix the abstract's overstatement.","headline":"Useful open BEST-RQ model release and SUPERB evaluation, but the Singapore-English gains need an explicit data-split statement before you trust them.","tokens_in":15859,"tokens_out":3201,"would_cite":true,"duration_ms":29267,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MERaLiON-SpeechEncoder is a 630M-parameter self-supervised speech encoder pretrained from scratch on 200,000 hours of audio that improves speech recognition on spontaneous and Singapore English speech while staying competitive across ten…","keywords":["self-supervised speech representation learning","BEST-RQ","speech foundation model","Singapore English","Singlish","masked language modeling","SUPERB benchmark","automatic speech recognition"],"falsifier":"Check the NSC evaluation split described in Appendix A.2 against the pretraining corpus in Table 1 by speaker ID or audio hash; if any test utterance appears in pretraining, re-run the Singapore English evaluation on a held-out corpus that was never in the pretraining mix and see whether the 15.2% average WER and the margin over WavLM large survive.","tokens_in":14822,"feed_emoji":"🎙️","tokens_out":11968,"duration_ms":89288,"temperature":0.7,"pith_summary":"MERaLiON-SpeechEncoder is a 630M-parameter speech foundation model pretrained from scratch on 200,000 hours of unlabelled audio with BEST-RQ, a BERT-style masked-language objective whose target labels come from frozen random projections rather than iterative clustering. The paper reports improved automatic speech recognition on spontaneous and Singapore English speech: 5.6% word error rate on TEDLIUMv3 test versus 6.2% for an in-house finetuned WavLM large, and a 15.2% average WER on the six-part Singapore English/Singlish National Speech Corpus subset after finetuning on only 420 hours. The same encoder reaches 2.1%/4.3% WER on LibriSpeech test-clean/test-other with 960 hours of finetuning, matching HuBERT large, and scores 82.62 on the ten-task SUPERB benchmark, between HuBERT large (82.25) and WavLM large (84.77). By releasing checkpoints and finetuning recipes, the authors aim to make a large-scale BEST-RQ encoder publicly available and useful for speech applications in Singapore and Southeast Asia.","feed_headline":"New 630M speech encoder beats WavLM on Singlish and spontaneous speech","feed_subtitle":"Pretrained on 200,000 hours, it cuts Singapore English WER while matching top encoders on ten SUPERB tasks.","key_machinery":"The load-bearing mechanism is BEST-RQ (BERT speech pre-training with random-projection quantizer), a masked-language objective that avoids the iterative k-means target computation used by HuBERT and WavLM. A frozen random projection matrix $A_j$ maps each downsampled, segment-level mean/variance-normalized Mel-spectrogram frame to one of 2048 random codewords $c_{ij}$ in each of 32 codebooks, and the label for a frame is $y^{\\mathrm{ref}}_j = \\arg\\min_i (A_j x - c_{ij})^\\top(A_j x - c_{ij})$. The 24-layer Conformer encoder is trained to predict these labels for masked frames using cross-entropy over the masked positions only. The paper's reported recipe changes—masking probability 0.4 instead of 0.01, Euclidean distance instead of cosine similarity, no L2 normalization, and 32 codebooks instead of one—are what make the random-projection targets effective at this scale.","core_discovery":"The central claim is that a BEST-RQ speech encoder trained at scale—200,000 hours, 630M parameters, 24 Conformer layers—yields representations that transfer especially well to the kinds of English that standard pretraining corpora under-represent: spontaneous talk and Singapore-accented English, including code-switched Singlish. Evidence for the claim is that finetuning on 420 hours of the National Speech Corpus produces a 15.2% average WER across the corpus's six parts, beating a Whisper large v3 finetuned on the full 8,169-hour cleaned corpus (16.9%) and outperforming WavLM large on the named-entity-heavy Part 2 and the spontaneous code-switching Part 4 by relative margins of 36.5% and 5.8%. On spontaneous TED talks, the encoder reaches 5.6% test WER, about 10% relative better than the finetuned WavLM baseline, and on LibriSpeech it ties HuBERT large at 2.1%/4.3%. The authors also show the encoder is not ASR-specific: its 82.62 SUPERB score places it between HuBERT large and WavLM large across phoneme, keyword, intent, slot, speaker, and emotion tasks.","pith_inferences":["Editorial inference: the report does not settle whether the NSC test audio was excluded from the 10,000 hours of NSC pretraining; if it was not, the Singapore English gains could partly reflect the encoder having heard the test speakers' unlabelled audio. This is testable by re-evaluating on a fresh Singapore English corpus.","Editorial inference: because the random-projection quantizer is language-agnostic and the pretraining mix already includes 30,000 hours of multilingual Common Voice, the same recipe should transfer to the planned Malay, Chinese, Tamil, Indonesian, Thai, and Vietnamese releases once comparable evaluation benchmarks exist.","Editorial inference: the 4x downsampling in the feature extractor may limit fine-grained time-sensitive tasks; the SUPERB speaker-verification and diarisation scores sit between HuBERT and WavLM, so future versions with less downsampling or pretraining augmentation could close the gap without losing the ASR gains."],"forward_implications":["Finetuning the encoder on 960 hours of LibriSpeech gives 2.1% WER on test-clean and 4.3% on test-other, matching HuBERT large and beating the in-house finetuned WavLM large (2.5%/4.6%).","On TEDLIUMv3 spontaneous speech, the encoder's 5.6% test WER is a relative improvement of about 10% over the in-house finetuned WavLM large's 6.2%.","On the Singapore English NSC subset, finetuning with 420 hours yields a 15.2% average WER across the six parts, better than Whisper large v3 finetuned on the full 8,169-hour corpus (16.9%) despite using roughly 5% of the finetuning data.","The SUPERB overall score of 82.62 indicates the encoder remains usable on non-ASR tasks such as phoneme recognition, keyword spotting, intent classification, speaker verification, and emotion recognition.","Public checkpoints and finetuning recipes let downstream developers adapt the encoder to their own tasks without repeating a 200,000-hour pretraining run."],"supporting_citations":[{"why":"defines the BEST-RQ objective and random-projection quantizer that the paper independently implements and scales.","marker":"Chiu et al. [2022]"},{"why":"introduces the multi-codebook BEST-RQ extension whose 16-codebook setup the paper validates and surpasses with 32 codebooks.","marker":"Zhang et al. [2023]"},{"why":"supplies the WavLM large model, its finetuning recipe, and the SUPERB overall-score formula used as baselines.","marker":"Chen et al. [2021b]"},{"why":"defines HuBERT's masked prediction of k-means targets, the main SSL comparison and a LibriSpeech baseline.","marker":"Hsu et al. [2021]"},{"why":"introduces wav2vec 2.0 contrastive SSL and the finetuning strategy the paper follows.","marker":"Baevski et al. [2020]"},{"why":"provides the National Speech Corpus, the Singapore English/Singlish data central to the paper's main claim.","marker":"Koh et al. [2019]"},{"why":"defines the SUPERB benchmark and the ten downstream tasks used to measure generality.","marker":"wen Yang et al. [2021]"},{"why":"provides the Whisper large zero-shot and finetuned ASR baselines compared on TEDLIUMv3 and NSC.","marker":"Radford et al. [2023]"},{"why":"supplies the Libri-light corpus used for the first 60K-hour pretraining phase and for baseline SSL models.","marker":"Kahn et al. [2020]"},{"why":"defines the Conformer architecture used for the 24-layer encoder stack.","marker":"Gulati et al. [2020]"}],"fun_headline_variants":["200k-hour speech encoder beats WavLM on Singlish and spontaneous speech","630M-param encoder improves Singapore English speech recognition","Speech foundation model for Singapore: beats WavLM on spontaneous talk","200k hours, 630M params: a speech encoder that improves Singlish"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Singapore English test utterances used to report word error rates were not among the audio clips the model saw, unlabelled, during pretraining; the report never states explicitly that this exclusion was made.","fun_headline_variants_meta":{"raw":{"variants":["200k-hour speech encoder beats WavLM on Singlish and spontaneous speech","630M-param encoder improves Singapore English speech recognition","Speech foundation model for Singapore: beats WavLM on spontaneous talk","200k hours, 630M params: a speech encoder that improves Singlish"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000838,"raw_usage":{"total_tokens":3674,"prompt_tokens":989,"completion_tokens":2685,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":2608}},"tokens_in":605,"tokens_out":2685,"duration_ms":17852,"temperature":1.0,"reasoning_tokens":2608,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:48:55.832427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the NSC evaluation split described in Appendix A.2 against the pretraining corpus in Table 1 by speaker ID or audio hash; if any test utterance appears in pretraining, re-run the Singapore English evaluation on a held-out corpus that was never in the pretraining mix and see whether the 15.2% average WER and the margin over WavLM large survive.","supporting_citations":[],"review_version":1}