{"id":"8019537c-5db8-41eb-bb11-ac6bc35ba03a","arxiv_id":"2506.07722","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new public benchmark for Arabic mispronunciation detection using Quranic recitation, with baseline models reaching F1 scores below 30%.","lead":"The paper introduces QuranMB.v1, a new test set of 98 Quranic verses read by 18 speakers with deliberately scripted pronunciation errors, and evaluates several self-supervised baseline models for Arabic mispronunciation detection. It is a first public benchmark for this task, but the test errors are generated from the authors' own confusion model, which limits external validity.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth errors in QuranMB.v1 are generated from the same confusion matrix used to synthesize the training data, so the reported F1=29.88 may not transfer to natural learner mispronunciations; the benchmark's external validity is untested.","rationale":"The paper's contribution is a first public benchmark for Arabic mispronunciation detection, and its headline result is the best baseline F1 of 29.88% on QuranMB.v1. For this contribution to be meaningful, the test set must measure what it claims to measure: mispronunciations that occur in real MSA pronunciation assessment. The reader identified the same core weakness: the test set's ground-truth errors are generated from the same confusion matrix used to create synthetic training data. My reading of the full text confirms this and adds a further aggravating detail: the test set errors are not collected from learners at all, but acted by native speakers following explicit instructions to produce modifications sampled from that matrix. This makes the measured performance a closed-loop evaluation of synthetic error patterns rather than an estimate of performance on natural learner errors. The paper's own Section 5 concedes that future work will need L2 speakers, which is an in-text admission that the current test set does not include the target population. This is not an internal inconsistency, but it is a serious external-validity problem for the 'unified benchmark' claim. The correct response is conditional acceptance: the resource release and baseline numbers are useful, but the benchmark's validity for real MSA mispronunciation detection must be demonstrated before the central claim is fully supported. Since the reader's verdict already reflects this conditionality, no verdict adjustment is needed.","tokens_in":7738,"tokens_out":3202,"duration_ms":39467,"concrete_test":"Collect a small held-out set of natural mispronunciations from L2 learners of MSA, e.g., 1-2 hours of expert-annotated recitation errors that are not sampled from the confusion matrix of reference [17], and evaluate the best model (mHuBERT, CMV-Ar+TTS) with unchanged hyperparameters. If the F1 on this natural set is substantially below 29.88, or if the error types fall largely outside the matrix, the test set is not representative of real mispronunciations. Additionally, verify that the HuggingFace repository is publicly downloadable, contains the confusion matrix, and provides the versioned test set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that QuranMB.v1 is a usable benchmark for MSA mispronunciation detection rests on the test set being representative of real pronunciation errors. Sections 2.2 and 2.4 show that both the synthetic training errors and the test-set errors are generated by randomly selecting characters or diacritics and replacing them using the same confusion matrix derived from reference [17]. The test set is additionally acted by native Arabic speakers instructed to produce those injected substitutions, not by L2 learners making natural errors. Consequently, the reported 29.88% F1 for mHuBERT on CMV-Ar+TTS is measured on a closed set of synthetic-error types that the model was explicitly trained to recognize; it does not estimate performance on naturally occurring mispronunciations or on error distributions outside the matrix. Section 5's stated future work of collecting L2 speaker data acknowledges this gap. The release promise also uses future tense in Sections 2.1 and 2.2, so the 'publicly available' status is not yet verifiable. These issues are addressable, but until the confusion matrix is validated against expert-annotated real learner errors and the data are actually released, the benchmark's external validity remains unestablished.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces QuranMB.v1, a benchmark test set for mispronunciation detection and diagnosis (MDD) in Modern Standard Arabic using Qur'anic recitation. It consists of 98 verses read by 18 native Arabic speakers who were instructed to produce specific pronunciation errors selected from a confusion matrix. The authors also build a 52-hour synthetic training corpus by modifying canonical vowelized transcripts with simulated errors and synthesizing speech with seven TTS voices. Several SSL-based baseline models (Wav2vec2, HuBERT, WavLM, mHuBERT) are evaluated under three training configurations, with the best reported result being an F1 of 29.88% for mHuBERT trained on combined real and synthetic data. The paper claims this is the first publicly available benchmark for Arabic mispronunciation detection in the Qur'anic recitation setting.","tokens_in":7934,"tokens_out":4880,"duration_ms":60304,"significance":"The contribution is timely and potentially useful: public MDD resources for Arabic are scarce, and the paper provides a documented pipeline spanning a specialized phoneme set, TTS-based error augmentation, a test set, and baseline evaluations. The release of QuranMB.v1, if actually public and properly validated, would give the community a controlled starting point for Arabic MDD research. However, the benchmark's external validity is currently unestablished because the test errors are generated from the same confusion matrix used to synthesize training data, and because no natural learner errors or expert annotations are involved. The baseline performance numbers should therefore be interpreted as measuring a model's fit to the authors' simulated error distribution, not its ability to detect natural Arabic mispronunciations.","major_comments":[{"comment":"Both the synthetic training errors (Section 2.2) and the test-set ground truth (Section 2.4) are generated by selecting characters or diacritics and replacing them using the same confusion matrix derived from the authors' prior work [17]. The test errors are additionally acted by native speakers instructed to produce those specific substitutions, so the test set does not contain independently observed learner mispronunciations. Consequently, the F1 scores in Table 2 (e.g., 29.88 for mHuBERT on CMV-Ar+TTS) measure agreement with the authors' simulated error model, not performance on naturally occurring errors. The manuscript should either validate the confusion matrix against expert-annotated real learner errors, use a held-out set of natural errors for testing, or explicitly limit the benchmark's claims to controlled synthetic-error scenarios.","section":"Sections 2.2 and 2.4"},{"comment":"The test set contains no naturally occurring mispronunciations: all errors are deliberately produced by native Arabic speakers following displayed instructions. There is no reported quality check of whether the speakers actually produced the intended substitutions, no inter-speaker consistency analysis, and no expert phonetician annotation of the recordings. This makes it difficult to know whether the ground-truth labels correspond to the acoustics in the audio. Please add quality-control statistics, such as agreement between intended and perceived errors or re-annotation of a subset by experts, and describe how the recorded speech was verified against the intended error patterns.","section":"Section 2.3"},{"comment":"The abstract and contributions claim that QuranMB.v1 is the 'first publicly available test set' and state that all models and datasets are available at the Hugging Face link, but the body repeatedly uses the future tense: the CMV-Ar corpus 'will be made publicly available' (Section 2.1), the TTS dataset 'will be publicly available' (Section 2.2), and the confusion dictionary 'will be publicly available' (Section 2.4). This inconsistency makes the central release claim unverifiable. Please provide stable identifiers (e.g., dataset card, DOI) and state the current accessibility, license, and access terms for QuranMB.v1 and the training corpora.","section":"Sections 2.1, 2.2, and 2.4"}],"minor_comments":[{"comment":"The sentence 'we randomly select four characters and/or diacritics and modify them based on a predefined confusion pairs matrix' is underspecified; please describe the exact sampling procedure, whether the number of modified tokens is fixed per transcript, and how the confusion matrix probabilities are applied, as this is essential for reproducibility.","section":"Section 2.2"},{"comment":"The definitions of Precision and Recall are nonstandard, and the roles of 'True Reject' and 'False Reject' should be clarified in the text. A brief justification that these quantities are computed only over the error-detection decision would help readers interpret Table 2.","section":"Section 3.4, Eq. (1)"},{"comment":"The table caption contains the misspelling 'mispronunications.' Additionally, the columns TA, FR, FA, CD, and ED are used in the caption without definitions; please define all abbreviations in the caption or in the text.","section":"Table 2"},{"comment":"The paper mentions that Arabic has 34 phonemes in the Introduction but then defines a phoneme inventory of 68 phonemes. Please clarify how the 68-phoneme set relates to the 34-phoneme description, particularly whether it includes geminated variants, emphatic variants, and positional allophones.","section":"Section 2.5"},{"comment":"The SSL model variants include English-only and multilingual models, but no Arabic-specific pre-trained model. If no suitable Arabic SSL model exists, state this explicitly; otherwise, its absence from the baselines should be justified.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The circularity concern raised by the stress-test note is valid and central: the test set and the synthetic training data are generated from the same confusion matrix, so the reported baselines do not demonstrate generalization to natural learner errors. This is addressable, and I do not recommend rejection if the authors add validation against expert-annotated or naturally occurring mispronunciations and clarify the public release status. Given that reference [17] is the authors' own prior work, the burden of demonstrating that its confusion matrix represents real errors is particularly high."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe headline: this is a genuinely useful resource paper for Arabic mispronunciation detection, but its central benchmark claim is weakened by a circularity in data construction. The paper introduces QuranMB.v1, the first public test set for MSA using Quranic recitation, plus a 68-phoneme inventory and baseline results. That is real value for a small subfield with no common evaluation set.\n\nWhat the paper does well: it documents the pipeline thoroughly, promises public releases on Hugging Face, and reports honest, modest baseline F1 scores (best 29.88%). The multilingual mHuBERT consistently beats monolingual SSL models, and the synthetic TTS data helps somewhat. These results are new and citable.\n\nThe soft spot is structural. Both the synthetic training data and the test set's ground-truth errors come from the same confusion matrix, derived from the authors' own prior work (ref [17]). Test speakers are instructed to realize those injected substitutions. So the 29.88% F1 is a measure of how well the model recognizes error types defined by that matrix—not necessarily how well it detects natural L2 mispronunciations. The paper flags this in Section 5 by listing L2 speaker collection as future work, so the circularity is a limitation rather than a fatal flaw. Still, 'unified benchmark' is over-strong until the confusion matrix is validated against expert-annotated real errors.\n\nTwo smaller issues: no human validation that speakers actually produced the intended errors, and no confidence intervals on baselines. The release promises use future tense for training sets, though the abstract claims availability; I'd check the HF repo before citing.\n\nOverall: this paper deserves a serious referee. The resource is valuable, the description is clear, and the concerns are addressable—validate the error model, add human annotation, publish with versioning. I'd send it to peer review with a request for revisions. For anyone in Arabic speech technology, the phoneme set and test set alone justify a read.","headline":"Useful first benchmark for Arabic MDD, but its test set is built from the same confusion matrix as the training data, so the reported F1 measures an in-distribution error model rather than real-world generalization.","tokens_in":8579,"tokens_out":5308,"would_cite":false,"duration_ms":53996,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces QuranMB.v1, the first publicly available benchmark for mispronunciation detection in Modern Standard Arabic using Qur'anic recitation, and reports that the strongest baseline—a multilingual HuBERT model trained on…","keywords":["Arabic pronunciation assessment","mispronunciation detection","Quranic recitation","Modern Standard Arabic","synthetic speech augmentation","confusion matrix","self-supervised speech models","benchmark dataset"],"falsifier":"Record a second test set in which Arabic speakers read the same verses freely without being told which errors to produce, have experts annotate the resulting mispronunciations, and run the same best model on it. If its F1 drops far below 29.88% or the distribution of error pairs diverges from the confusion matrix, the benchmark's claim to represent realistic MSA mispronunciations would be falsified.","tokens_in":7499,"feed_emoji":"🎙️","tokens_out":8674,"duration_ms":99477,"temperature":0.7,"pith_summary":"The authors set out to give Arabic pronunciation-assessment research a common yardstick by releasing QuranMB.v1, a public test benchmark built from 98 Qur'anic verses read by 18 native Arabic speakers with deliberately injected pronunciation errors. They accompany it with a full pipeline: a 68-phoneme inventory tailored to Modern Standard Arabic, an in-the-wild real-speech training set, and 52 hours of synthetic speech in which errors are inserted by a confusion matrix and rendered by text-to-speech. Using frozen self-supervised speech encoders with CTC decoding, they report that the best baseline, mHuBERT trained on real plus synthetic data, reaches 29.88% F1 on mispronunciation detection. The headline point is not that the problem is solved but that a standardized resource now exists on which future approaches can be compared, and the low score quantifies how much room remains.","feed_headline":"Best model scores only 29.88% F1 on new Arabic recitation test","feed_subtitle":"The public QuranMB.v1 benchmark sets a yardstick for Arabic mispronunciation detection and shows the task is unsolved.","key_machinery":"The load-bearing object is QuranMB.v1, the test benchmark itself, together with the error-generation protocol used to build it: a confusion matrix derived from phoneme-similarity data maps each Arabic character or diacritic to likely mispronounced counterparts, including deletions. The same matrix is used to corrupt canonical transcripts before TTS rendering, yielding the 52-hour synthetic training corpus; this symmetry is what makes the controlled errors fully annotated. The recognition pipeline is a frozen self-supervised encoder whose layer-weighted features feed a two-layer Bi-LSTM with CTC loss, with greedy decoding producing phoneme sequences that are compared against phonetizer-derived targets. The evaluation follows the standard TA/TR/FA/FR and CD/ED categorization, with F1 on the reject class as the headline metric.","core_discovery":"The paper's central claim is that Qur'anic recitation, read in Modern Standard Arabic without tajweed constraints, is a workable case study for benchmarking Arabic mispronunciation detection, and that a publicly released test set plus a reproducible training pipeline can support fair comparisons. The authors assert that their synthetic TTS corpus, generated by randomly modifying four characters or diacritics per canonical transcript according to a confusion matrix derived from phoneme similarity data, is competitive with wild-collected speech for training MDD models: mHuBERT trained on the combined corpus attains the best F1 (29.88%), and even TTS-only training beats several English-only SSL baselines trained on real speech. They also claim the 68-phoneme inventory, which merges emphatic-context vowel variants into single phonemes and marks gemination by doubled symbols, is appropriate for the task.","pith_inferences":["Because the same confusion matrix appears to generate both the synthetic training errors and the scripted test errors, the reported F1 may be optimistic relative to natural, unscripted mispronunciations; a benchmark containing spontaneously occurring errors would be needed to test that transfer.","The benchmark deliberately reads MSA without Tajweed rules, so it measures segmental errors such as consonant substitutions rather than prosodic recitation mistakes; researchers applying it to recitation quality should not expect it to cover those.","The protocol of cueing speakers with highlighted modified text likely produces non-spontaneous, carefully timed errors, which could make the test either easier or harder than natural errors in ways the current numbers do not reveal.","Extending the same confusion-matrix and TTS pipeline to second-language learners would require enlarging the phoneme inventory with non-Arabic sounds, as the authors themselves note for future work."],"forward_implications":["Future Arabic mispronunciation detection systems can be compared on a stable public test set instead of private, ad hoc evaluations.","Because synthetic TTS-only training already rivals real-speech-only training for some baselines, controlled synthetic mispronunciation data is a credible route around annotation scarcity.","Multilingual self-supervised pretraining transfers to Arabic better than English-only pretraining in these experiments, so subsequent CAPT work should start from multilingual encoders.","The reported F1 below 30% means current models falsely accept most mispronunciations, so data curation and specialized modeling, not just larger pretraining, are the near-term levers.","The 68-phoneme inventory and the evaluation protocol become reusable infrastructure for other Arabic pronunciation tasks, including future extensions to Tajweed-oriented checking."],"supporting_citations":[{"why":"Supplies the real Modern Standard Arabic training speech (CMV-Ar) used in all real-data baselines.","marker":"[15]"},{"why":"Supplies the rationale that TTS-synthesized mispronounced speech can substitute for scarce natural error data.","marker":"[16]"},{"why":"Supplies the phoneme-similarity confusion matrix used to inject errors into transcripts for both TTS training data and test-set prompts.","marker":"[17]"},{"why":"Supplies the Arabic phonetizer that converts vowelized transcripts into the 68-phoneme target sequences.","marker":"[18]"},{"why":"Supplies the frozen-encoder training recipe, with weighted layer features and a task head, used for all baseline systems.","marker":"[19]"},{"why":"Supplies the CTC loss and greedy decoding used to turn speech into phoneme sequences.","marker":"[20]"},{"why":"Supplies the multilingual HuBERT model whose representations drive the best-performing baseline.","marker":"[23]"},{"why":"Establishes the TA/TR/FA/FR evaluation categories and the F1 definition used for the headline result.","marker":"[25]"}],"fun_headline_variants":["QuranMB.v1: first public benchmark for Arabic mispronunciation","Best model hits just 29.88% F1 on Quranic recitation test","Arabic mispronunciation detection: new benchmark reveals hard task","New public benchmark for Quranic recitation stumps all models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole evaluation rests on a table of which Arabic sounds are commonly swapped for which; if that table does not reflect how real speakers actually mispronounce Modern Standard Arabic, both the synthetic training data and the scripted test errors are unrealistic.","fun_headline_variants_meta":{"raw":{"variants":["QuranMB.v1: first public benchmark for Arabic mispronunciation","Best model hits just 29.88% F1 on Quranic recitation test","Arabic mispronunciation detection: new benchmark reveals hard task","New public benchmark for Quranic recitation stumps all models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000504,"raw_usage":{"total_tokens":2414,"prompt_tokens":852,"completion_tokens":1562,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":1485}},"tokens_in":468,"tokens_out":1562,"duration_ms":12693,"temperature":1.0,"reasoning_tokens":1485,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:27:06.457622+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a second test set in which Arabic speakers read the same verses freely without being told which errors to produce, have experts annotate the resulting mispronunciations, and run the same best model on it. If its F1 drops far below 29.88% or the distribution of error pairs diverges from the confusion matrix, the benchmark's claim to represent realistic MSA mispronunciations would be falsified.","supporting_citations":[{"cited_title":"Computer-assisted pronunciation train- ing—speech synthesis is almost all you need,","cited_arxiv_id":null,"evidence_quote":"Supplies the rationale that TTS-synthesized mispronounced speech can substitute for scarce natural error data."},{"cited_title":"wav2vec 2.0: A framework for self-supervised learning of speech representations,","cited_arxiv_id":null,"evidence_quote":"Supplies the phoneme-similarity confusion matrix used to inject errors into transcripts for both TTS training data and test-set prompts."},{"cited_title":"This phonetizer was optimized for phonetic coverage in speech synthesis","cited_arxiv_id":null,"evidence_quote":"Supplies the Arabic phonetizer that converts vowelized transcripts into the 68-phoneme target sequences."},{"cited_title":"The mgb-2 challenge: Arabic multi- dialect broadcast media recognition,","cited_arxiv_id":null,"evidence_quote":"Supplies the frozen-encoder training recipe, with weighted layer features and a task head, used for all baseline systems."},{"cited_title":"To obtain the phoneme sequence during inference, CTC greedy decoding is used","cited_arxiv_id":null,"evidence_quote":"Supplies the CTC loss and greedy decoding used to turn speech into phoneme sequences."},{"cited_title":"Development of automated tajweed checking system for children in learning quran,","cited_arxiv_id":null,"evidence_quote":"Supplies the multilingual HuBERT model whose representations drive the best-performing baseline."},{"cited_title":"SpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation","cited_arxiv_id":"2211.00923","evidence_quote":"Establishes the TA/TR/FA/FR evaluation categories and the F1 definition used for the headline result."}],"review_version":1}