Pith. sign in

REVIEW 3 major objections 4 minor 92 references

NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read NADI 2025 establishes a three-task benchmark for spoken Arabic dialects and reports the first leaderboard scores for dialect identification, speech recognition, and diacritic restoration.

desk verdict Solid benchmark report that deserves minor revisions: the new diacritic-restoration shared task and blind test sets are real contributions, but the 'first' claim is oversized and the test-label quality control is unverified. read the letter →

arxiv 2509.02038 v2 pith:UHEFKN5Y submitted 2025-09-02 cs.CL cs.SD

classification cs.CLcs.SD
keywords Arabicdialectidentificationautomaticspeechrecognitiondiacriticrestorationsharedtaskmultidialectalcode-switchingbenchmarkNLP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents NADI 2025, the sixth edition of the Nuanced Arabic Dialect Identification shared task and, by the authors' account, the first shared task devoted to multidialectal Arabic speech processing. It defines three subtasks—spoken dialect identification across eight country-level varieties, automatic speech recognition of those varieties including code-switched speech, and diacritic restoration for MSA, Classical Arabic, dialects, and code-switched text—and supplies new blind test sets, an adaptation set, baselines, and evaluation metrics. Eight teams submitted 100 valid test-phase entries. The reported best scores are 79.8% dialect-identification accuracy with an average cost of 17.88, 35.68/12.20 WER/CER on ASR, and 55/13 WER/CER on diacritic restoration. A sympathetic reading is that the shared task creates a common benchmark and credible reference points for an under-resourced area; the numbers also show large headroom, especially for ASR.

What carries the argument

The load-bearing mechanism is the evaluation protocol itself: three newly curated blind test sets, a deliberately small adaptation set, and standardized metrics—accuracy and the LRE 2022 average-cost measure for dialect ID, and a normalized WER/CER pipeline for ASR and diacritic restoration. The blind test sets let any system be measured on the same speech, while the adaptation set is meant to force transfer-learning and adaptation strategies rather than training from scratch. The baselines give the leaderboard meaning by providing the score a system must beat.

What would settle it

Take a random sample of the blind test utterances, have a second set of fluent speakers independently assign dialect labels and diacritics, and measure agreement with the official references; low agreement (for example, below 90% on dialect labels or below 0.8 kappa on diacritics) would mean the leaderboard scores are not a stable measure of system quality.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is that a speech-focused Arabic dialect benchmark is feasible and that it produces useful, comparable measurements. The organizers curated blind test sets—eight hours of speech across eight dialects for dialect ID, 10,807 utterances for ASR, and 1,332 manually diacritized utterances for diacritic restoration—together with a small adaptation set drawn from different series than the test set to reduce speaker overlap. Baselines include a fine-tuned speaker-recognition model for dialect ID, a zero-shot large pretrained speech model for ASR, and text-only, ASR-based, and multimodal models for diacritics. The best submitted systems beat these baselines on a

Load-bearing premise

The blind test sets carry manual dialect labels and diacritic references that are assumed to be correct and consistent; the paper reports fluent-speaker verification and manual annotation but no agreement or quality statistics.

Editorial extensions

If this is right

  • Future dialect-identification systems can be compared directly against 79.8% accuracy on the same eight-way blind test set.
  • The ASR leaderboard, with best average WER 35.68, becomes the reference point for progress on multidialectal and code-switched Arabic speech.
  • The 55 WER baseline for diacritic restoration gives the first public yardstick for dialectal and code-switched diacritization, where none previously existed.
  • The winning ASR system's large gain over the zero-shot baseline shows that weakly supervised pretraining on large unverified Arabic audio plus continued supervised fine-tuning is a productive recipe for low-resource dialectal ASR.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the adaptation set is small and the test set is blind, the dialect-ID leaderboard is as much a measure of data efficiency and transfer as of ceiling accuracy; scores could shift if participants were allowed to train on all available external dialect speech.
  • Since WER compares against a single reference transcript, the reported ASR scores may underestimate systems that produce valid but different transcriptions; an evaluation with multiple accepted references would likely lower WER.
  • The three-task structure invites a single end-to-end model that identifies dialect, transcribes, and restores diacritics jointly; the paper points toward unified dialect-aware speech technology but does not build it.
  • Country-level labels are coarse proxies for a continuous dialect landscape, so high scores on these eight labels should not be read as mastery of Arabic dialects generally; this is a paper-acknowledged limitation with practical consequences for deployment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports on the sixth edition of the NADI shared task, presented as the first multidialectal Arabic speech processing shared task. It introduces three subtasks: spoken Arabic dialect identification (8 country-level dialects), multidialectal Arabic ASR, and diacritic restoration for spoken Arabic varieties, along with new curated datasets, blind test sets, and evaluation protocols. Forty-four teams registered; eight teams made 100 valid test-phase submissions. The best reported results are 79.8% accuracy (Cavg 17.88) for Subtask 1, 35.68/12.20 WER/CER for Subtask 2, and 55/13 WER/CER for Subtask 3. The paper describes baselines, participant systems, and limitations, and releases datasets and leaderboards.

Significance. If the blind-test labels and evaluation protocol are reliable, this paper provides a valuable benchmark for Arabic dialect speech processing: it introduces new test sets, documents baselines, and reports results from independent third-party systems. The release of datasets and leaderboards, the use of multiple complementary subtasks, and the candid acknowledgment of WER/CER limitations with single references are all strengths. The main risk to the central claim is that test-label quality is not demonstrated: no inter-annotator agreement or quality-control analysis is reported for the manually verified/annotated blind test sets, so the leaderboard numbers and the relative system rankings should be treated with some caution until that evidence is supplied.

major comments (3)
  1. [Section 3.1 and Section 3.3] The benchmark's validity depends on the correctness of the blind test labels, but no reliability evidence is reported. For Subtask 1, dialects are said to be 'verified by fluent speakers' (Section 3.1), yet there is no information about the number of annotators, annotation instructions, agreement, or adjudication. For Subtask 3, Table 1 and Section 3.3 state that test sets were created by 'manually annotating random subsets,' again without any reliability metric. If the test labels are noisy, the reported accuracy/WER/CER values and the relative rankings could be distorted. Please add an annotation protocol, inter-annotator agreement (e.g., Cohen's kappa or pairwise agreement), and a brief error/label-noise analysis, or at least clearly justify why label noise is unlikely to affect the reported conclusions.
  2. [Section 4.3, Table 4] There is an internal inconsistency: the text states that Munsit achieves the lowest overall average WER/CER of 35.68/12.10, while the abstract and Table 4 report 35.68/12.20. This must be reconciled. More broadly, the paper reports no confidence intervals or significance tests, although several rankings are close (e.g., Subtask 1: 79.8 vs. 76.4; Subtask 2: 35.68 vs. 38.53 WER). Without uncertainty estimates, the reader cannot tell whether the reported ordering reflects meaningful system differences. Please correct the CER value and add significance tests or otherwise qualify the leaderboard comparisons.
  3. [Section 3.1] The adaptation/test split is described as designed to 'minimize the influence of potentially overlapping speakers' and to avoid series overlap, but speaker overlap is not guaranteed. If the same speaker appears in both the adaptation set and the blind test set, Subtask 1 results could partly reflect speaker matching rather than dialect identification. Please report an overlap analysis (e.g., speaker-level statistics if speaker IDs are available), or explain why the series-level disjointness is sufficient to prevent speaker-level leakage. This is directly relevant to the validity of the ADI leaderboard.
minor comments (4)
  1. [Abstract] There is a LaTeX artifact in the participant distribution: 'five teams{\ae}' should read 'five teams'.
  2. [Section 4.2 / Table 5] The baseline numbering is inconsistent: the text lists baselines (I) text-only CATT, (II) ASR-based ArTST, and (III) multi-modal, while Table 5 uses 'Baseline-I (ASR based)', 'Baseline-II (text-only)', and 'Baseline-II (multi-modal)'. Please align the numbering/labels.
  3. [Section 4.3 / Table 4] Several typos appear ('preformernce' in Section 4.3, 'diaritize' in Table 1, 'inference inference' in Section 5.3). Also, WER values above 100 for the baseline and MarsadLab in Table 4 are not explained; a brief note that WER can exceed 100 when insertions dominate would help.
  4. [Section 4.3] The statement that Munsit 'obtain the lowest overall average WER/CER scores' is followed by per-dialect exceptions (Moroccan, Mauritanian) that are described but not shown in a consistent way. Consider presenting the per-dialect rankings in a clearer format so the reader can verify the aggregate claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: NADI 2025 reports independent third-party system results evaluated on newly created test sets; self-citations are provenance, not load-bearing premises.

full rationale

NADI 2025 is a shared-task organization paper rather than a predictive derivation. The headline numbers (79.8% accuracy, 35.68/12.20 WER/CER, 55/13 WER/CER) are produced by independent participating teams on blind test sets curated by the organizers; they are not fitted parameters renamed as predictions, nor are they entailed by any equation in the paper. Self-citations to prior NADI editions and the Casablanca corpus serve as provenance for dataset construction, evaluation normalization, and baseline design; they do not force the reported outcomes. The claim of being 'the first multidialectal Arabic speech processing shared task' is a historical/descriptive claim, not a result derived from an input that assumes it. The Limitations section does acknowledge that WER/CER may be misleading for dialectal speech with multiple valid references, and the paper does not report inter-annotator agreement for its manually verified/annotated test labels; these are validity and reliability concerns rather than circularity under the defined patterns, because no claimed result reduces by construction to its own input or to a self-citation chain. The benchmark results are externally produced and falsifiable, so no load-bearing circular step is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper is an empirical benchmark report; it introduces datasets and results but no new theoretical entities. No free parameters are fitted, no new entities are postulated.

assumptions (3)
  • domain assumption The blind test sets are correctly labeled (dialects, diacritics)
    Results depend on annotation quality; no inter-annotator agreement is reported (Sections 3.1, 3.3).
  • domain assumption Adaptation and test sets do not share speakers
    The paper says it 'aims to minimize' overlap (Section 3.1), not guarantee it.
  • domain assumption The normalization pipeline (removing diacritics, normalizing Hamza, etc.) is applied consistently to outputs and references
    Inconsistent application would distort WER/CER (Section 3.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task." pith.science (2026). https://pith.science/paper/UHEFKN5Y

@misc{pith2026250902038,
  author       = {Pith},
  title        = {Pith review of: NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UHEFKN5Y}},
  note         = {Machine review of arXiv:2509.02038}
}
read the original abstract

We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken dialect identification (Subtask 1), speech recognition (Subtask 2), and diacritic restoration for spoken dialects (Subtask 3). A total of 44 teams registered, and during the testing phase, 100 valid submissions were received from eight unique teams. The distribution was as follows: 34 submissions for Subtask 1 "five teams{\ae}, 47 submissions for Subtask 2 "six teams", and 19 submissions for Subtask 3 "two teams". The best-performing systems achieved 79.8% accuracy on Subtask 1, 35.68/12.20 WER/CER (overall average) on Subtask 2, and 55/13 WER/CER on Subtask 3. These results highlight the ongoing challenges of Arabic dialect speech processing, particularly in dialect identification, recognition, and diacritic restoration. We also summarize the methods adopted by participating teams and briefly outline directions for future editions of NADI.

Figures

Figures reproduced from arXiv: 2509.02038 by the authors.

Figure 1
Figure 1. Overview of the NADI 2025 shared tasks. 2018; Darwish et al., 2021; Abdul-Mageed et al., 2020, 2023). At the same time, many downstream applications—from automatic transcription and vir￾tual assistants to text-to-speech and educational tools—depend on accurate handling of dialectal speech and the diacritics that indicate short vow￾els and phonological features. Existing systems trained on CA/MSA (Elmadany et al., 20… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

92 extracted references · 68 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Ahmed Amine Ben Abdallah, Ata Kabboudi, Amir Kanoun, and Salah Zaiem. 2023. https://arxiv.org/abs/2309.11327 Leveraging data collection and unsupervised learning for code-switched tunisian arabic automatic speech recognition . Preprint, arXiv:2309.11327

  4. [4]

    Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan, and Kareem Darwish. 2021. Qadi: Arabic dialect identification in the wild. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 1--10, Kyiv, Ukraine (Virtual). Association for Computational Linguistics

  5. [5]

    Muhammad Abdul-Mageed, Hassan Alhuzali, and Mohamed Elaraby. 2018. You tweet what you speak: A city-level dataset of arabic dialects. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA)

  6. [6]

    Muhammad Abdul-Mageed, AbdelRahim Elmadany, Chiyu Zhang, El Moatez Billah Nagoudi, Houda Bouamor, and Nizar Habash. 2023. Nadi 2023: The fourth nuanced arabic dialect identification shared task. arXiv preprint arXiv:2310.16117

  7. [7]

    Muhammad Abdul-Mageed, Amr Keleg, AbdelRahim Elmadany, Chiyu Zhang, Injy Hamed, Walid Magdy, Houda Bouamor, and Nizar Habash. 2024. Nadi 2024: The fifth nuanced arabic dialect identification shared task. arXiv preprint arXiv:2407.04910

  8. [8]

    Muhammad Abdul - Mageed, Chiyu Zhang, Houda Bouamor, and Nizar Habash. 2020. https://www.aclweb.org/anthology/2020.wanlp-1.9/ NADI 2020: The first nuanced arabic dialect identification shared task . In Proceedings of the Fifth Arabic Natural Language Processing Workshop, WANLP@COLING 2020, Barcelona, Spain (Online), December 12, 2020, pages 97--110. Assoc...

Show all 92 references
  1. [9]

    Muhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim Elmadany, Houda Bouamor, and Nizar Habash. 2022. Nadi 2022: The third nuanced arabic dialect identification shared task. arXiv preprint arXiv:2210.09582

  2. [10]

    Muhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim Elmadany, and Lyle Ungar. 2020. Toward micro-dialect identification in diaglossic and code-switched environments. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5855--5876,...

  3. [11]

    Elmadany, Houda Bouamor, and Nizar Habash

    Muhammad Abdul - Mageed, Chiyu Zhang, AbdelRahim A. Elmadany, Houda Bouamor, and Nizar Habash. 2021. https://www.aclweb.org/anthology/2021.wanlp-1.28/ NADI 2021: The second nuanced arabic dialect identification shared task . In Proceedings of the Sixth Arabic Natural Language ...

  4. [12]

    Badr Abdullah, Yusser Al-Ghussein, Zena Al-Khalili, Ömer Özyilmaz, Matias Valdenegro-Toro, Simon Ostermann, and Dietrich Klakow. 2025. Saarland-groningen at nadi 2025 shared task: Effective dialectal arabic speech processing under data constraints. In The Third Arabic Natural ...

  5. [13]

    Kathrein Abu Kwaik, Motaz Saad, Stergios Chatzikyriakidis, and Simon Dobnik. 2018. Shami: A corpus of levantine arabic dialects. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)

  6. [14]

    Maryam Khalifa Al Ali and Hanan Aldarmaki. 2024. https://aclanthology.org/2024.sigul-1.26 Mixat: A data set of bilingual emirati- E nglish speech . In Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages @ LREC-COLING 2024, pages 222...

  7. [15]

    Abandah, Adham Alsharkawi, and Maha Dawas

    Mohammad Al-Fetyani, Muhammad Al-Barham, Gheith A. Abandah, Adham Alsharkawi, and Maha Dawas. 2022. https://doi.org/10.1109/SLT54892.2023.10022652 MASC : Massive arabic speech corpus . In Proceedings of the 2022 IEEE Spoken Language Technology Workshop (SLT), page 1002206

  8. [16]

    Rania Al-Sabbagh and Roxana Girju. 2012. Yadac: Yet another dialectal arabic corpus. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 2882--2889, Istanbul, Turkey. European Language Resources Association (ELRA)

  9. [17]

    Nora Al-Twairesh, Rawan N. Al-Matham, Nora Madi, Nada Almugren, Al-Hanouf Al-Aljmi, Shahad Alshalan, Raghad Alshalan, Nafla Alrumayyan, Shams Al-Manea, Sumayah Bawazeer, Nourah Al-Mutlaq, Nada Almanea, Waad Bin Huwaymil, Dalal Alqusair, Reem Alotaibi, Suha Al-Senaydi, and Abee...

  10. [18]

    Faris Alasmary, Orjuwan Zaafarani, and Ahmad Ghannam. 2024. https://arxiv.org/abs/2407.03236 Catt: Character-based arabic tashkeel transformer . Preprint, arXiv:2407.03236

  11. [19]

    Sanad ALBawwab and Omar Qawasmeh. 2025. Lahjati at nadi 2025 a ecapa-wavlm fusion with multi-stage optimization. In The Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou. Association for Computational Linguistics

  12. [20]

    Hanan Aldarmaki and Ahmad Ghannam. 2023. Diacritic recognition performance in arabic asr. In Proc. Interspeech 2023, pages 361--365

  13. [21]

    Ahmed Ali, Peter Bell, James Glass, Yacine Messaoui, Hamdy Mubarak, Steve Renals, and Yifan Zhang. 2016. https://doi.org/10.1109/SLT.2016.7846277 The mgb-2 challenge: Arabic multi-dialect broadcast media recognition . In 2016 IEEE Spoken Language Technology Workshop (SLT), pag...

  14. [22]

    Ahmed Ali, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James Glass, Steve Renals, and Khalid Choukri. 2019. The mgb-5 challenge: Recognition and dialect identification of dialectal arabic speech. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop...

  15. [23]

    Ahmed Ali, Stephan Vogel, and Steve Renals. 2017. Speech recognition challenge in the wild: Arabic mgb-3. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 316--322. IEEE

  16. [24]

    Khalid Almeman, Mark Lee, and Ali Abdulrahman Almiman. 2013. https://doi.org/10.1109/ICCSPA.2013.6487288 Multi dialect arabic speech parallel corpora . In 2013 1st International Conference on Communications, Signal Processing, and their Applications (ICCSPA), pages 1--6

  17. [25]

    Israa Alsarsour, Esraa Mohamed, Reem Suwaileh, and Tamer Elsayed. 2018. Dart: A large dataset of dialectal arabic tweets. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)

  18. [26]

    Althobaiti

    Maha J. Althobaiti. 2022. Creation of annotated country-level dialectal arabic resources: An unsupervised approach. Natural Language Engineering, 28(5):607--648

  19. [27]

    Tyers, and Gregor Weber

    Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M. Tyers, and Gregor Weber. 2020. Common voice: A massively-multilingual speech corpus. In Proceedings of the 12th Conference on Language Resources...

  20. [28]

    Rafiul Biswas, Shimaa Ibrahim, Mabrouka Bessghaier, Firoj Alam, and Wajdi Zaghouani

    Kais Attia, Md. Rafiul Biswas, Shimaa Ibrahim, Mabrouka Bessghaier, Firoj Alam, and Wajdi Zaghouani. 2025. Marsadlab at nadi: Arabic dialect identification and speech recognition using ecapa-tdnn and whisper. In The Third Arabic Natural Language Processing Conference (ArabicNL...

  21. [29]

    Matthew Baas, Benjamin van Niekerk , and Herman Kamper. 2023. https://doi.org/10.21437/Interspeech.2023-419 Voice conversion with just nearest neighbors . In Interspeech 2023, pages 2053--2057

  22. [30]

    As-Said Muh \'a mmad Badawi. 1973. Mustawayat al-arabiyya al-muasira fi Misr. Dar al-maarif

  23. [31]

    Nurpeiis Baimukan, Houda Bouamor, and Nizar Habash. 2022. Hierarchical aggregation of dialectal data for arabic dialect identification. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (LREC 2022), pages 4586--4596, Marseille, France. European Lang...

  24. [32]

    Lo \" c Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, and 1 others. 2023. Seamless: Multilingual expressive and streaming speech translation. arXiv preprint arXiv:2312.05187

  25. [33]

    Houda Bouamor, Nizar Habash, and Kemal Oflazer. 2014. A multidialectal parallel corpus of arabic. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 1240--1245, Reykjavik, Iceland. European Language Resources Association (ELRA)

  26. [34]

    Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, and Kemal Oflazer. 2018. http://www.lrec-conf.org/proceedings/lrec2018/summaries/351.html The MADAR arabic dialect corpus...

  27. [35]

    Kristen Brustad. 2000. The Syntax of Spoken Arabic: A Comparative Study of Moroccan, Egyptian, Syrian, and Kuwaiti Dialects. Georgetown University Press

  28. [36]

    Shammur A Chowdhury, Ahmed Ali, Suwon Shon, and James R Glass. 2020. What does an end-to-end dialect identification model learn about non-dialectal information? In INTERSPEECH, pages 462--466

  29. [37]

    Al-Natsheh, Houda Bouamor, Karim Bouzoubaa, Violetta Cavalli-Sforza, Samhaa R

    Kareem Darwish, Nizar Habash, Mourad Abbas, Hend Al-Khalifa, Huseein T. Al-Natsheh, Houda Bouamor, Karim Bouzoubaa, Violetta Cavalli-Sforza, Samhaa R. El-Beltagy, Wassim El-Hajj, Mustafa Jarrar, and Hamdy Mubarak. 2021. https://doi.org/10.1145/3447735 A P anoramic survey of N ...

  30. [38]

    Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020. https://doi.org/10.21437/Interspeech.2020-2650 Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification . In Interspeech 2020, pages 3830--3834

  31. [39]

    Mona Diab, Nizar Habash, Owen Rambow, Mohamed Altantawy, and Yassine Benajiba. 2010. Colaba: Arabic dialect annotation and processing. In Proceedings of the LREC Workshop on Semitic Language Processing, pages 66--74

  32. [40]

    Amirbek Djanibekov, Hawau Olamide Toyin, Raghad Alshalan, Abdullah Alitr, and Hanan Aldarmaki. 2025. https://arxiv.org/abs/2411.05872 Dialectal coverage and generalization in arabic speech recognition . Preprint, arXiv:2411.05872

  33. [41]

    Mahmoud El-Haj. 2020. Habibi -- a multi dialect multi national arabic song lyrics corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 1318--1326, Marseille, France. European Language Resources Association (ELRA)

  34. [42]

    Heba Elfardy, Mohamed Al-Badrashiny, and Mona Diab. 2014. Aida: Identifying code switching in informal arabic text. In Proceedings of the First Workshop on Computational Approaches to Code Switching, pages 94--101, Doha, Qatar. Association for Computational Linguistics

  35. [43]

    Salman Elgamal, Ossama Obeid, Mhd Kabbani, Go Inoue, and Nizar Habash. 2024. https://doi.org/10.18653/v1/2024.acl-long.792 A rabic diacritics in the wild: Exploiting opportunities for improved diacritization . In Proceedings of the 62nd Annual Meeting of the Association for Co...

  36. [44]

    Haroun Elleuch, Salima Mdhaffar, Yannick Estève, and Fethi Bougares. 2025 . https://doi.org/ 10.21437/Interspeech.2025-884 ADI-20: Arabic Dialect Identification dataset and models . In Interspeech 2025 , pages 2775--2779

  37. [45]

    Haroun Elleuch, Youssef Saidi, Salima Mdhaffar, Yannick Estève, and Fethi Bougares. 2025. Elyadata & lia at nadi 2025: Asr and adi subtasks. In The Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou. Association for Computational Linguistics

  38. [46]

    AbdelRahim Elmadany, Muhammad Abdul-Mageed, and 1 others. 2022. Arat5: Text-to-text transformers for arabic language generation. In Proceedings of the 60th annual meeting of the association for computational linguistics (Volume 1: Long papers), pages 628--647

  39. [47]

    Abdelrahim Elmadany, Muhammad Abdul-Mageed, and 1 others. 2023 a . Octopus: A multitask model and toolkit for arabic natural language generation. In Proceedings of ArabicNLP 2023, pages 232--243

  40. [48]

    AbdelRahim Elmadany, ElMoatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023 b . Orca: A challenging benchmark for arabic language understanding. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9559--9586, Toronto, Canada. Association for Computa...

  41. [49]

    Mohamed Lotfy Elrefai. 2025. Unicorn at nadi 2025 subtask 3: Gemm3n-dr: Audio-text diacritic restoration via fine-tuned multimodal arabic llm. In The Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou. Association for Computational Linguistics

  42. [50]

    Ali Fadel, Ibraheem Tuffaha, Bara ' Al-Jawarneh, and Mahmoud Al-Ayyoub. 2019. https://doi.org/10.18653/v1/D19-5229 Neural A rabic text diacritization: State of the art results and a novel approach for machine translation . In Proceedings of the 6th Workshop on Asian Translatio...

  43. [51]

    Hassan Gadalla, Hanaa Kilany, Howaida Arram, Ashraf Yacoub, Alaa El-Habashi, Amr Shalaby, Krisjanis Karins, Everett Rowson, Robert MacIntyre, Paul Kingsbury, David Graff, and Cynthia McLemore. 1997. Callhome egyptian arabic transcripts ldc97t19. Web Download

  44. [52]

    Ahmad Ghannam, Naif Alharthi, Faris Alasmary, Kholood Al Tabash, Shouq Sadah, and Lahouari Ghouti. 2025. Abjad ai at nadi 2025: Catt-whisper: Multimodal diacritic restoration using text and speech representations. In The Third Arabic Natural Language Processing Conference (Ara...

  45. [53]

    Alex Graves, Santiago Fern \'a ndez, Faustino Gomez, and J \"u rgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning, pages 369--376

  46. [54]

    Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and 1 others. 2020. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100

  47. [55]

    Karim El Haff, Mustafa Jarrar, Tymaa Hammouda, and Fadi Zaraket. 2022. https://aclanthology.org/2022.lrec-1.82.pdf C urras + B aladi: T owards a L evantine C orpus . In Proceedings of the International Conference on Language Resources and Evaluation(LREC 2022), Marseille, France

  48. [56]

    Nagham Hamad, Mohammed Khalilia, and Mustafa Jarrar. 2025. http://www.jarrar.info/publications/HKJ25.pdf K onooz: M ulti-domain M ulti-dialect C orpus for N amed E ntity R ecognition . In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, ...

  49. [57]

    Injy Hamed, Ngoc Thang Vu, and Slim Abdennadher. 2020. https://aclanthology.org/2020.lrec-1.523/ A rz E n: A speech corpus for code-switched E gyptian A rabic- E nglish . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4237--4246, Marseille, F...

  50. [58]

    Harrat, M

    S. Harrat, M. Abbas, K. Meftouh, and K. Smaili. 2013. https://doi.org/10.21437/Interspeech.2013-373 Diacritics restoration for arabic dialect texts . In Interspeech 2013, pages 1429--1433

  51. [59]

    Salima Harrat, Karima Meftouh, Mourad Abbas, and Kamel Sma \"i li. 2014. Building resources for algerian arabic dialects. In INTERSPEECH 2014, 15th Annual Conference of the International Speech Communication Association, pages 2123--2127, Singapore. ISCA

  52. [60]

    Richard S. Harrell. 1962. A Short Reference Grammar of Moroccan Arabic: With Audio CD. Georgetown Classics in Arabic Language and Linguistics. Georgetown University Press

  53. [61]

    Clive Holes. 2004. Modern Arabic: Structures, Functions, and Varieties. Georgetown Classics in Arabic Language and Linguistics. Georgetown University Press

  54. [62]

    Elsayed Issa, Mohammed AlShakhori, Reda AlBahrani, and Gus Hahn-Powell. 2021. Country-level arabic dialect identification using rnns with and without linguistic features. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 276--281, Kyiv, Ukraine (Vi...

  55. [63]

    Mustafa Jarrar, Nizar Habash, Faeq Alrimawi, Diyam Akra, and Nasser Zalmout. 2016. Curras: An annotated corpus for the palestinian arabic dialect. Language Resources and Evaluation, pages 1--31

  56. [64]

    Mustafa Jarrar, Fadi Zaraket, Tymaa Hammouda, Daanish Masood Alavi, and Martin Waahlisch. 2023. https://doi.org/1 10.1109/AICCSA59173.2023.10479250 L isan: Y emeni, I rqi, L ibyan, and S udanese A rabic D ialect C opora with M orphological A nnotations . In The 20th IEEE/ACS I...

  57. [65]

    Salam Khalifa, Nizar Habash, Dana Abdulrahim, and Sara Hassan. 2016. A large scale corpus of gulf arabic. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4282--4289, Portorož, Slovenia. European Language Resources Asso...

  58. [66]

    Taku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226

  59. [67]

    Ajinkya Kulkarni, Atharva Kulkarni, Sara Abedalmon'em Mohammad Shatnawi, and Hanan Aldarmaki. 2023. https://doi.org/10.21437/Interspeech.2023-2224 Clartts: An open-source classical arabic text-to-speech corpus . In 2023 INTERSPEECH, pages 5511--5515

  60. [68]

    Mingfei Lau, Qian Chen, Yeming Fang, Tingting Xu, Tongzhou Chen, and Pavel Golik. 2025. https://doi.org/10.18653/v1/2025.acl-long.370 Data quality issues in multilingual speech datasets: The need for sociolinguistic awareness and proactive language planning . In Proceedings of...

  61. [69]

    Yooyoung Lee, Craig Greenberg, Lisa Mason, and Elliot Singer. 2022. Nist 2022 language recognition evaluation plan

  62. [70]

    Shervin Malmasi, Marcos Zampieri, Nikola Ljube s i \'c , Preslav Nakov, Ahmed Ali, and J \"o rg Tiedemann. 2016. Discriminating between similar languages and A rabic dialect identification: A report on the third DSL shared task. In Proceedings of the third workshop on NLP for ...

  63. [71]

    Karima Meftouh, Salima Harrat, Salma Jamoussi, Mourad Abbas, and Kamel Smaili. 2015. Machine translation experiments on padic: A parallel arabic dialect corpus. In Proceedings of the 29th Pacific Asia Conference on Language, Information and Computation, pages 26--34, Shanghai, China

  64. [72]

    Hamdy Mubarak and Kareem Darwish. 2014. Using twitter to collect a multi-dialectal corpus of arabic. In Proceedings of the EMNLP 2014 Workshop on Arabic Natural Language Processing (ANLP), pages 1--7, Doha, Qatar. Association for Computational Linguistics

  65. [73]

    El Moatez Billah Nagoudi, Ahmed El-Shangiti, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2023. Dolphin: A challenging and diverse benchmark for arabic nlg. arXiv preprint arXiv:2305.14989

  66. [74]

    Amal Nayouf, Mustafa Jarrar, Fadi zaraket, Tymaa Hammouda, and Mohamad-Bassam Kurdy. 2023. https://doi.org/10.18653/v1/2023.arabicnlp-1.2 N âbra: S yrian A rabic D ialects with M orphological A nnotations . In Proceedings of the 1st Arabic Natural Language Processing Conferenc...

  67. [75]

    Ossama Obeid, Mohammad Salameh, Houda Bouamor, and Nizar Habash. 2019. Adida: Automatic dialect identification for arabic. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 6--11, Minnea...

  68. [76]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. https://doi.org/10.48550/ARXIV.2212.04356 Robust speech recognition via large-scale weak supervision . arXiv preprint

  69. [77]

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492--28518. PMLR

  70. [78]

    Fatiha Sadat, Farzindar Kazemi, and Atefeh Farzindar. 2014. Automatic identification of arabic language varieties and dialects in social media. In Proceedings of the Second Workshop on Natural Language Processing for Social Media (SocialNLP), pages 22--27, Dublin, Ireland. Ass...

  71. [79]

    Mohammad Salameh, Houda Bouamor, and Nizar Habash. 2018. Fine-grained arabic dialect identification. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1332--1344, Santa Fe, New Mexico, USA. Association for Computational Linguistics

  72. [80]

    Mahmoud Salhab, Marwan Elghitany, Shameed Sait, Syed Sibghat Ullah, Mohammad Abusheikh, and Hasan Abusheikh. 2025 a . Advancing arabic speech recognition through large-scale weakly supervised learning. arXiv preprint arXiv:2504.12254

  73. [81]

    Mahmoud Salhab, Shameed Sait, Mohammad Abusheikh, and Hasan Abusheikh. 2025 b . Munsit at nadi 2025 shared task 2: Pushing the boundaries of multidialectal arabic asr with weakly supervised pretraining and continual supervised fine-tuning. In The Third Arabic Natural Language ...

  74. [82]

    Sara Shatnawi, Sawsan Alqahtani, and Hanan Aldarmaki. 2024. Automatic restoration of diacritics for speech data sets. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Lo...

  75. [83]

    Suwon Shon, Ahmed Ali, Younes Samih, Hamdy Mubarak, and James Glass. 2020. Adi17: A fine-grained arabic dialect identification dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8244--8248. IEEE

  76. [84]

    Peter Sullivan, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2023. https://doi.org/10.21437/Interspeech.2023-1005 On the robustness of arabic speech dialect identification . In Interspeech 2023, pages 5326--5330

  77. [85]

    Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, and 1 others. 2024. Casablanca: Data and models for multidialectal arabic speech recognition. arXiv ...

  78. [86]

    Hawau Toyin, Amirbek Djanibekov, Ajinkya Kulkarni, and Hanan Aldarmaki. 2023. Artst: Arabic text and speech transformer. In Proceedings of ArabicNLP 2023, pages 41--51

  79. [87]

    Magdy, and Hanan Aldarmaki

    Hawau Toyin, Rufael Marew, Humaid Alblooshi, Samar M. Magdy, and Hanan Aldarmaki. 2025 . https://doi.org/ 10.21437/Interspeech.2025-1550 ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis . In Interspeech 2025 , pages 4808--4812

  80. [88]

    o rgen Valk and Tanel Alum \

    J \"o rgen Valk and Tanel Alum \"a e. 2021. Voxlingua107: a dataset for spoken language recognition. In 2021 IEEE Spoken Language Technology Workshop (SLT), pages 652--658. IEEE

  81. [89]

    Abdul Waheed, Bashar Talafha, Peter Sullivan, Abdelrahim Elmadany, and Muhammad Abdul-Mageed. 2023. Voxarabica: A robust dialect-aware arabic speech recognition system. In Proceedings of ArabicNLP 2023, pages 441--449

  82. [90]

    Wajdi Zaghouani and Anis Charfi. 2018. Arap-tweet: A large multi-dialect twitter corpus for gender, age and language variety identification. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Lang...

  83. [91]

    Zaidan and Chris Callison-Burch

    Omar F. Zaidan and Chris Callison-Burch. 2011. The arabic online commentary dataset: An annotated dataset of informal arabic with high dialectal content. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pa...

  84. [92]

    Taha Zerrouki and Amar Balla. 2017. https://doi.org/10.1016/j.dib.2017.01.011 Tashkeela: Novel corpus of arabic vocalized texts, data for auto-diacritization systems . Data in Brief, 11

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.