{"id":"0e48ca9c-9c6e-4e4d-aa58-d88dbb535b1a","arxiv_id":"2501.01034","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The authors release MNSC, the largest standardized multitask spoken Singlish corpus, and SingAudioLLM, a multimodal model that sets strong baselines on ASR, spoken QA, dialogue summarization, and paralinguistic QA.","lead":"This paper builds a new multitask speech dataset for Singlish, a Singaporean English creole, and trains a multimodal audio model that can transcribe, answer questions about, summarize, and analyze speakers' voices. The resource is meant to give researchers a standardized way to study spoken Singlish and to test audio AI models on it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10–30% SOTA claim is not statistically secured: SQA/SDS results in Table 2 come from 100-sample test sets scored by an undisclosed model-as-judge, with no error bars or significance tests.","rationale":"The paper's dataset construction is a genuine contribution: it reorganizes the National Speech Corpus into standardized multitask splits, adds human-verified test sets, and releases code, model, and data. Synthetic LLM-generated training targets are a lesser concern in themselves, because training-label noise can be tolerable when the evaluation set is reliable; the load-bearing issue is evaluation reliability. The reader's weakest_assumption points in the same direction, but I would sharpen it: the problem is not primarily that the training data are synthetic, but that the SQA/SDS test sets are small and scored by an undisclosed model-as-judge, so the reported margins may not be real. A concrete statistical check—paired bootstrap confidence intervals on the 100 test samples plus a second human rater—would settle this. If the margins survive, the SOTA claim is credible; if not, the paper should present the model as competitive and report intervals. This concern does not change the reader's CONDITIONAL verdict, but it identifies the precise condition the revision must satisfy.","tokens_in":14668,"tokens_out":5016,"duration_ms":49844,"concrete_test":"Rescore every SQA/SDS test response in Table 2 with two independent human annotators (or with a fully disclosed judge and fixed rubric), then compute paired bootstrap 95% confidence intervals for each SingAudioLLM-versus-strongest-baseline difference over the 100 test items. If the interval includes zero for a majority of the bolded SQA/SDS cells, or if the ranking flips under the second annotator, the 10–30% SOTA claim should be replaced by a task-by-task comparison with intervals reported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that SingAudioLLM achieves state-of-the-art performance, outperforming prior models by 10–30%—rests heavily on the SQA and SDS scores in Table 2. Section 3.2 reports only 100 human-annotated test samples per SQA/SDS subtask, and Section 5 says evaluation uses 'WER or model-as-judge scores' without naming the judge, its prompt, or agreement with human raters. With n=100, the standard error of a proportion is up to 5 percentage points, so many reported leads over the cascade or Qwen2-Audio-Instruct baselines are within plausible noise: SQA Part 4 is 51.8 vs 53.8 (SingAudioLLM behind), SQA Part 6 is 64.0 vs 64.0 (tied), and SDS Part 5 is 50.0 vs 49.0 (a 1-point lead). Without error bars, paired significance tests, or rater agreement, the claimed margins are not distinguishable from noise, and the abstract's '10–30%' overstates what the table establishes. Table 3 is an additional caution: FT-Whisper beats SingAudioLLM on 4 of 6 MNSC ASR cells, so blanket 'state-of-the-art' phrasing is broader than the paper's own evidence. The dataset resource is likely useful, but the performance headline is the least secure part of the argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MNSC, a multitask spoken Singlish benchmark derived from the National Speech Corpus, with standardized splits and human-annotated test sets for ASR, spoken question answering (SQA), spoken dialogue summarization (SDS), and paralinguistic question answering (PQA). The authors also propose SingAudioLLM, a fusion-style AudioLLM that combines a Whisper encoder, an adaptor, and an LLM decoder trained jointly on all four tasks. They report strong in-domain results, claim a 10-30% improvement over prior AudioLLMs and cascaded baselines, and release the dataset, code, and model. The main empirical evidence is in Tables 2 and 3, with component analyses in Figure 3.","tokens_in":15031,"tokens_out":3576,"duration_ms":32339,"significance":"If the benchmark is sound and the reported comparisons are statistically reliable, MNSC would be a valuable resource for a genuinely underserved language, and the release of standardized splits, human-annotated test sets, and code would lower the barrier for future work. The paper also has useful component analyses (encoder size, adaptor choice, decoder size) that speak to practical design questions for low-resource AudioLLMs. However, the central performance claim is not yet supported: the SQA/SDS results rest on 100-sample test sets with undisclosed model-as-judge scoring and no error bars, and several of the paper's own tables show baseline models beating SingAudioLLM on specific tasks. The dataset contribution is promising; the state-of-the-art claim needs substantially stronger evidence.","major_comments":[{"comment":"The SQA/SDS test sets contain only n=100 human-annotated samples per subtask, and the paper reports no error bars, confidence intervals, or paired significance tests. For an accuracy score with n=100, the standard error can reach roughly 5 percentage points, so most of the reported margins are within plausible noise: SingAudioLLM is behind the cascade on MNSC-SQA PART 4 (51.8 vs 53.8), tied on MNSC-SQA PART 6 (64.0 vs 64.0), and leads by only 1 point on MNSC-SDS PART 5 (50.0 vs 49.0). The abstract's claim of outperforming prior models by 10-30% is not supportable from Table 2 without confidence intervals or paired tests.","section":"§3.2, Table 2"},{"comment":"The evaluation section states that metrics use 'WER or model-as-judge scores' but never names the judge model, the prompt template, or the scoring protocol, and no agreement with human ratings is reported. This is load-bearing for the SQA/SDS results because the training targets were synthesized by LLAMA-3.1-70B (§3.2); if the same model family is used as the judge, the comparison may be biased toward outputs that resemble that model's style. Please disclose the judge, report human-model agreement, and provide an error analysis on the 100-sample test sets.","section":"§5, evaluation protocol"},{"comment":"The blanket 'state-of-the-art' claim is not supported by the paper's own evidence. In Table 3, FT-Whisper achieves lower WER than SingAudioLLM on MNSC-ASR Parts 2, 3, and 4, and on SEAME-Dev-SGE and SEAME-Dev-MAN; in Table 2, the cascade model beats SingAudioLLM on MNSC-SQA Part 4 and ties on Part 6. The claim should be restricted to the settings where it actually holds (for example, certain AudioLLM comparisons on MNSC), or the authors should provide an aggregate result with uncertainty that justifies the headline.","section":"Tables 2 and 3, §5"},{"comment":"The dataset construction caps dialogue duration at 30 seconds, excludes recordings that 'could not be perfectly aligned,' and the Limitations section admits that a substantial portion of the source corpus was excluded. Since no statistics or sensitivity analyses are reported for these exclusions, readers cannot tell whether the MNSC benchmark and the resulting model rankings are representative of the full NSC or are artifacts of the curation thresholds. Please report the amount and characteristics of excluded data and, if feasible, test robustness to the 30-second cap and the alignment threshold.","section":"§3.1 and Limitations"}],"minor_comments":[{"comment":"There are typos and inconsistencies: 'models's' in the Abstract, 'Coprus' in §3.1, 'Pralinguistic' in Table 1, and 'NMSC' instead of 'MNSC' in the Table 3 caption. Please proofread the manuscript.","section":"Abstract, §3.1, Table 1"},{"comment":"The header 'MNSC-SQS-PART 3-6' should be 'MNSC-SQA-PART 3-6', and the appendix inconsistently uses 'Sing-AudioLLM' instead of 'SingAudioLLM'. Please unify the naming.","section":"Appendix A, Table 4"},{"comment":"The legend and axes do not clearly separate the three studies (encoder size, adaptor, decoder); please mark the different conditions consistently and add confidence intervals or at least multiple-seed markers to support the qualitative claims about performance differences.","section":"Figure 3"},{"comment":"Appendix B is titled 'Hardward' instead of 'Hardware'. In addition, the paper does not state how many random seeds were used for LoRA training or for decoding, which is needed for reproducibility of the numerical comparisons.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The dataset resource is likely a meaningful contribution, but the performance headline is not yet empirically secure. The authors should be asked to provide significance testing, disclose the model-as-judge details, and revise the state-of-the-art wording to match what the tables actually show. If they can add these analyses, the paper may be suitable for publication. I do not see a fundamental flaw in the resource itself, so I am not recommending rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the MNSC dataset is the real contribution here, and it should enter the literature. The SingAudioLLM SOTA claim is not supported by the paper's own numbers, so treat that part with skepticism.\n\nWhat's genuinely new: MNSC is the first standardized multitask benchmark for spoken Singlish, derived from the National Speech Corpus. You get ASR, spoken QA, dialogue summarization, and paralinguistic QA with consistent splits and human-verified test sets. For a low-resource Creole, that's a useful public resource. I also appreciate the ablation study on encoders, adaptors, and decoder sizes—it gives practical guidance for people building AudioLLMs.\n\nThe soft spot is the model claim. The abstract says SOTA and '10-30%' over prior models, but Table 2 tells a more measured story. On SQA, the cascade beats SingAudioLLM on PART4 (53.8 vs 51.8) and ties on PART6. In Table 3, fine-tuned Whisper is better on four of six ASR cells. So the SOTA phrase is only true if you ignore the strongest baselines. The margins that do exist are also fragile: the SQA/SDS test sets are 100 samples each, so the standard error is around 5 points. A 1-2 point lead is not meaningful, and the paper reports no error bars or significance tests. On top of that, the model-as-judge evaluation for SQA/SDS is never specified—no judge model, no prompt, no human agreement. That's a reproducibility hole, and it matters because the training targets are LLM-generated, so an LLM judge could be biased toward the model's style.\n\nTo be fair, none of this sinks the dataset. The resource is valuable regardless of which model wins, and the human-verified test sets are a solid basis for future benchmarking. But the performance chapter needs to be rewritten: report confidence intervals, disclose the judge, and stop calling the model SOTA when the tables don't support it.\n\nI'd send this to peer review. With revisions on the evaluation transparency and the claim calibration, it's a publishable resource paper. Definitely worth having a serious referee look at it.","headline":"The MNSC dataset is a real contribution; the SingAudioLLM SOTA claim does not survive its own tables.","tokens_in":15557,"tokens_out":2874,"would_cite":true,"duration_ms":26301,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces the largest standardized spoken Singlish corpus and a multitask audio-LLM that reports 10–30% gains over prior systems.","keywords":["Singlish","spoken language understanding","audio large language models","multimodal large language models","speech recognition","spoken question answering","dialogue summarization","paralinguistic analysis"],"falsifier":"The quickest test would be to take the released MNSC test sets and compute confidence intervals on the reported scores; with 100 items per subtask, if the 10–30% margins overlap between SingAudioLLM and the next-best model, the claimed lead is not established. A stronger test would be an independent, larger set of Singlish clips (e.g., 500 per task) with human-verified labels, scored by the same models: if the ordering or margin changes, the central claim fails.","tokens_in":14498,"feed_emoji":"🗣️","tokens_out":11424,"duration_ms":86167,"temperature":0.7,"pith_summary":"Singlish is a widely spoken creole in Singapore, but its spoken form has lacked the standardized, multitask resources that modern speech models need. This paper tries to close that gap by releasing MNSC, a cleaned and re-split version of the largest available Singlish speech corpus, with training data for speech recognition, spoken question answering, dialogue summarization, and paralinguistic question answering, plus human-verified test sets. It also proposes SingAudioLLM, a single audio-language model trained jointly on all four tasks. The paper reports that SingAudioLLM outperforms prior audio-LLMs and a cascade baseline by 10–30% on these benchmarks, and that the corpus helps even a standard ASR model adapt to Singlish. If the claims hold, researchers gain a reproducible benchmark and a reference model for a low-resource creole that has so far been limited mostly to ASR.","feed_headline":"New Singlish speech corpus and model beat prior audio-LLMs by 10–30%","feed_subtitle":"A standardized four-task benchmark lets one model transcribe, answer, summarize, and identify accents in Singlish.","key_machinery":"The central object is MNSC, the Multitask National Speech Corpus: standardized splits over the National Speech Corpus, with synthetic training data for SQA and SDS created by prompting an LLM with transcripts, and human-annotated test sets. The central mechanism is the fusion pipeline of SingAudioLLM: a Whisper encoder turns audio into features, a Conv-1D adapter downsamples and projects them into the text-token space, and the Gemma2-9B LLM decoder, tuned with LoRA, reads the audio tokens alongside a task instruction and generates the answer. The paper argues that jointly training this pipeline across ASR, SQA, SDS, and PQA lets the model share acoustic and semantic information, while keeping the decoder's instruction-following ability intact.","core_discovery":"On its own terms, the paper's central claim is that spoken Singlish understanding can be lifted from a single-task ASR problem to a four-task multimodal benchmark, and that an end-to-end audio LLM trained jointly on those tasks is the right way to attack it. To do this, the paper constructs MNSC by taking the National Speech Corpus, cleaning alignment errors, standardizing train/test splits, adding LLM-synthesized question-answer pairs and dialogue summaries for training, and building human-annotated test sets for SQA and SDS. SingAudioLLM uses a Whisper-large-v3 encoder, a Conv-1D adapter that compresses audio features, and a Gemma2-9B-Instruct decoder trained with LoRA; the same model answers prompts for ASR, SQA, SDS, and PQA. On MNSC it reports the best scores in all four task groups, and it claims a 10–30% improvement over other AudioLLMs and a Whisper-to-LLM cascade, with the largest gaps in paralinguistic tasks such as accent identification.","pith_inferences":["If the synthetic training recipe transfers, the same pipeline—clean an existing speech corpus, let an LLM generate QA and summaries from transcripts, and verify a small test set by humans—could be applied to other under-resourced creoles and dialects, not just Singlish.","The PQA results are suggestive rather than conclusive: accent labels come from speaker metadata, so a model might be exploiting lexical cues rather than acoustic ones; a follow-up with de-lexicalized or masked audio would separate those.","Because SQA and SDS test sets contain only 100 samples per subtask, the 10–30% margins may not be statistically stable; reporting confidence intervals or expanding the test set would turn a promising result into a solid benchmark.","The observation that a task-specific fine-tuned Whisper beats the multitask model on ASR hints at a specialization–generalization trade-off; routing or an expert mixture could let one model keep both the multitask strengths and the ASR edge."],"forward_implications":["The MNSC splits give the community a common yardstick for Singlish spoken ASR, SQA, SDS, and PQA, replacing the previous situation where benchmark splits did not exist.","A single jointly trained audio-LLM can match or beat a cascade that transcribes first and then runs an LLM, at least on the MNSC test sets, so end-to-end fusion is a viable design for creole speech understanding.","Fine-tuning Whisper on MNSC-ASR data cuts Singlish word error rates relative to the off-the-shelf Whisper and transfers to the SEAME code-switching dataset, indicating the corpus has value beyond the proposed model.","Decoder size matters for reasoning-heavy tasks such as QA and summarization, but not much for ASR, while encoder strength matters across tasks; this points to where future Singlish speech models should spend parameters.","The release of datasets, model, and code turns Singlish from a language with only ASR resources into a testbed for multilingual and code-switched multimodal understanding."],"supporting_citations":[{"why":"Supplies the raw National Speech Corpus audio and transcriptions that MNSC cleans, re-splits, and extends to four tasks.","marker":"Koh et al., 2019"},{"why":"Provides the Whisper-large-v3 encoder used in SingAudioLLM and the strong ASR baseline against which adaptation is measured.","marker":"Radford et al., 2023"},{"why":"The Llama-3.1-70B-Instruct model is prompted with transcripts to generate the SQA and SDS training pairs, so the multitask training data depends on it.","marker":"Dubey et al., 2024"},{"why":"Supplies the Gemma2-9B-Instruct decoder that SingAudioLLM uses as its language model backbone for the main results.","marker":"Team et al., 2024"},{"why":"Defines Qwen2-Audio-Instruct, the strongest prior AudioLLM baseline compared in the main results.","marker":"Chu et al., 2024"},{"why":"Provides the SEAME Mandarin-English code-switching corpus used to test zero-shot transfer of Singlish-trained models.","marker":"Lyu et al., 2010"},{"why":"Supplies the SALMONN baseline and the window-based Q-Former adapter design that the paper adapts and compares with Conv-1D.","marker":"Tang et al., 2024"},{"why":"Supplies LoRA, the low-rank adaptation method used to tune the LLM decoder while preserving its instruction-following abilities.","marker":"Hu et al., 2021"}],"fun_headline_variants":["SingAudioLLM sets new SotA on spoken Singlish, 10-30% above prior audio-LLMs","Four-task spoken Singlish benchmark: MNSC corpus + SingAudioLLM beat others by 10-30%","One model for ASR, QA, summarization, accent ID in Singlish: 10-30% gains","SingAudioLLM: end-to-end audio LLM for 4 Singlish tasks, 10-30% better than cascades"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's value and the model's reported lead rest on machine-generated QA and summary training data being a faithful stand-in for real spoken comprehension, and on 100 human-checked test items per subtask being enough to separate models.","fun_headline_variants_meta":{"raw":{"variants":["SingAudioLLM sets new SotA on spoken Singlish, 10-30% above prior audio-LLMs","Four-task spoken Singlish benchmark: MNSC corpus + SingAudioLLM beat others by 10-30%","One model for ASR, QA, summarization, accent ID in Singlish: 10-30% gains","SingAudioLLM: end-to-end audio LLM for 4 Singlish tasks, 10-30% better than cascades"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00088,"raw_usage":{"total_tokens":3802,"prompt_tokens":945,"completion_tokens":2857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":2736}},"tokens_in":561,"tokens_out":2857,"duration_ms":18427,"temperature":1.0,"reasoning_tokens":2736,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:36:10.159972+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The quickest test would be to take the released MNSC test sets and compute confidence intervals on the reported scores; with 100 items per subtask, if the 10–30% margins overlap between SingAudioLLM and the next-best model, the claimed lead is not established. A stronger test would be an independent, larger set of Singlish clips (e.g., 500 per task) with human-verified labels, scored by the same models: if the ordering or margin changes, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the raw National Speech Corpus audio and transcriptions that MNSC cleans, re-splits, and extends to four tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the SEAME Mandarin-English code-switching corpus used to test zero-shot transfer of Singlish-trained models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SALMONN baseline and the window-based Q-Former adapter design that the paper adapts and compares with Conv-1D."}],"review_version":1}