REVIEW 3 major objections 4 minor 92 references
NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read NADI 2025 establishes a three-task benchmark for spoken Arabic dialects and reports the first leaderboard scores for dialect identification, speech recognition, and diacritic restoration.
desk verdict Solid benchmark report that deserves minor revisions: the new diacritic-restoration shared task and blind test sets are real contributions, but the 'first' claim is oversized and the test-label quality control is unverified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the evaluation protocol itself: three newly curated blind test sets, a deliberately small adaptation set, and standardized metrics—accuracy and the LRE 2022 average-cost measure for dialect ID, and a normalized WER/CER pipeline for ASR and diacritic restoration. The blind test sets let any system be measured on the same speech, while the adaptation set is meant to force transfer-learning and adaptation strategies rather than training from scratch. The baselines give the leaderboard meaning by providing the score a system must beat.
What would settle it
Take a random sample of the blind test utterances, have a second set of fluent speakers independently assign dialect labels and diacritics, and measure agreement with the official references; low agreement (for example, below 90% on dialect labels or below 0.8 kappa on diacritics) would mean the leaderboard scores are not a stable measure of system quality.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that a speech-focused Arabic dialect benchmark is feasible and that it produces useful, comparable measurements. The organizers curated blind test sets—eight hours of speech across eight dialects for dialect ID, 10,807 utterances for ASR, and 1,332 manually diacritized utterances for diacritic restoration—together with a small adaptation set drawn from different series than the test set to reduce speaker overlap. Baselines include a fine-tuned speaker-recognition model for dialect ID, a zero-shot large pretrained speech model for ASR, and text-only, ASR-based, and multimodal models for diacritics. The best submitted systems beat these baselines on a
Load-bearing premise
The blind test sets carry manual dialect labels and diacritic references that are assumed to be correct and consistent; the paper reports fluent-speaker verification and manual annotation but no agreement or quality statistics.
Editorial extensions
If this is right
- Future dialect-identification systems can be compared directly against 79.8% accuracy on the same eight-way blind test set.
- The ASR leaderboard, with best average WER 35.68, becomes the reference point for progress on multidialectal and code-switched Arabic speech.
- The 55 WER baseline for diacritic restoration gives the first public yardstick for dialectal and code-switched diacritization, where none previously existed.
- The winning ASR system's large gain over the zero-shot baseline shows that weakly supervised pretraining on large unverified Arabic audio plus continued supervised fine-tuning is a productive recipe for low-resource dialectal ASR.
Reading between the lines
- Because the adaptation set is small and the test set is blind, the dialect-ID leaderboard is as much a measure of data efficiency and transfer as of ceiling accuracy; scores could shift if participants were allowed to train on all available external dialect speech.
- Since WER compares against a single reference transcript, the reported ASR scores may underestimate systems that produce valid but different transcriptions; an evaluation with multiple accepted references would likely lower WER.
- The three-task structure invites a single end-to-end model that identifies dialect, transcribes, and restores diacritics jointly; the paper points toward unified dialect-aware speech technology but does not build it.
- Country-level labels are coarse proxies for a continuous dialect landscape, so high scores on these eight labels should not be read as mastery of Arabic dialects generally; this is a paper-acknowledged limitation with practical consequences for deployment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports on the sixth edition of the NADI shared task, presented as the first multidialectal Arabic speech processing shared task. It introduces three subtasks: spoken Arabic dialect identification (8 country-level dialects), multidialectal Arabic ASR, and diacritic restoration for spoken Arabic varieties, along with new curated datasets, blind test sets, and evaluation protocols. Forty-four teams registered; eight teams made 100 valid test-phase submissions. The best reported results are 79.8% accuracy (Cavg 17.88) for Subtask 1, 35.68/12.20 WER/CER for Subtask 2, and 55/13 WER/CER for Subtask 3. The paper describes baselines, participant systems, and limitations, and releases datasets and leaderboards.
Significance. If the blind-test labels and evaluation protocol are reliable, this paper provides a valuable benchmark for Arabic dialect speech processing: it introduces new test sets, documents baselines, and reports results from independent third-party systems. The release of datasets and leaderboards, the use of multiple complementary subtasks, and the candid acknowledgment of WER/CER limitations with single references are all strengths. The main risk to the central claim is that test-label quality is not demonstrated: no inter-annotator agreement or quality-control analysis is reported for the manually verified/annotated blind test sets, so the leaderboard numbers and the relative system rankings should be treated with some caution until that evidence is supplied.
major comments (3)
- [Section 3.1 and Section 3.3] The benchmark's validity depends on the correctness of the blind test labels, but no reliability evidence is reported. For Subtask 1, dialects are said to be 'verified by fluent speakers' (Section 3.1), yet there is no information about the number of annotators, annotation instructions, agreement, or adjudication. For Subtask 3, Table 1 and Section 3.3 state that test sets were created by 'manually annotating random subsets,' again without any reliability metric. If the test labels are noisy, the reported accuracy/WER/CER values and the relative rankings could be distorted. Please add an annotation protocol, inter-annotator agreement (e.g., Cohen's kappa or pairwise agreement), and a brief error/label-noise analysis, or at least clearly justify why label noise is unlikely to affect the reported conclusions.
- [Section 4.3, Table 4] There is an internal inconsistency: the text states that Munsit achieves the lowest overall average WER/CER of 35.68/12.10, while the abstract and Table 4 report 35.68/12.20. This must be reconciled. More broadly, the paper reports no confidence intervals or significance tests, although several rankings are close (e.g., Subtask 1: 79.8 vs. 76.4; Subtask 2: 35.68 vs. 38.53 WER). Without uncertainty estimates, the reader cannot tell whether the reported ordering reflects meaningful system differences. Please correct the CER value and add significance tests or otherwise qualify the leaderboard comparisons.
- [Section 3.1] The adaptation/test split is described as designed to 'minimize the influence of potentially overlapping speakers' and to avoid series overlap, but speaker overlap is not guaranteed. If the same speaker appears in both the adaptation set and the blind test set, Subtask 1 results could partly reflect speaker matching rather than dialect identification. Please report an overlap analysis (e.g., speaker-level statistics if speaker IDs are available), or explain why the series-level disjointness is sufficient to prevent speaker-level leakage. This is directly relevant to the validity of the ADI leaderboard.
minor comments (4)
- [Abstract] There is a LaTeX artifact in the participant distribution: 'five teams{\ae}' should read 'five teams'.
- [Section 4.2 / Table 5] The baseline numbering is inconsistent: the text lists baselines (I) text-only CATT, (II) ASR-based ArTST, and (III) multi-modal, while Table 5 uses 'Baseline-I (ASR based)', 'Baseline-II (text-only)', and 'Baseline-II (multi-modal)'. Please align the numbering/labels.
- [Section 4.3 / Table 4] Several typos appear ('preformernce' in Section 4.3, 'diaritize' in Table 1, 'inference inference' in Section 5.3). Also, WER values above 100 for the baseline and MarsadLab in Table 4 are not explained; a brief note that WER can exceed 100 when insertions dominate would help.
- [Section 4.3] The statement that Munsit 'obtain the lowest overall average WER/CER scores' is followed by per-dialect exceptions (Moroccan, Mauritanian) that are described but not shown in a consistent way. Consider presenting the per-dialect rankings in a clearer format so the reader can verify the aggregate claim.
Circularity Check
No circular derivation: NADI 2025 reports independent third-party system results evaluated on newly created test sets; self-citations are provenance, not load-bearing premises.
full rationale
NADI 2025 is a shared-task organization paper rather than a predictive derivation. The headline numbers (79.8% accuracy, 35.68/12.20 WER/CER, 55/13 WER/CER) are produced by independent participating teams on blind test sets curated by the organizers; they are not fitted parameters renamed as predictions, nor are they entailed by any equation in the paper. Self-citations to prior NADI editions and the Casablanca corpus serve as provenance for dataset construction, evaluation normalization, and baseline design; they do not force the reported outcomes. The claim of being 'the first multidialectal Arabic speech processing shared task' is a historical/descriptive claim, not a result derived from an input that assumes it. The Limitations section does acknowledge that WER/CER may be misleading for dialectal speech with multiple valid references, and the paper does not report inter-annotator agreement for its manually verified/annotated test labels; these are validity and reliability concerns rather than circularity under the defined patterns, because no claimed result reduces by construction to its own input or to a self-citation chain. The benchmark results are externally produced and falsifiable, so no load-bearing circular step is present.
Assumptions & free parameters
assumptions (3)
- domain assumption The blind test sets are correctly labeled (dialects, diacritics)
- domain assumption Adaptation and test sets do not share speakers
- domain assumption The normalization pipeline (removing diacritics, normalizing Hamza, etc.) is applied consistently to outputs and references
Cite this review
Pith. "Pith review of NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task." pith.science (2026). https://pith.science/paper/UHEFKN5Y
@misc{pith2026250902038,
author = {Pith},
title = {Pith review of: NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task},
year = {2026},
howpublished = {\url{https://pith.science/paper/UHEFKN5Y}},
note = {Machine review of arXiv:2509.02038}
}
read the original abstract
We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken dialect identification (Subtask 1), speech recognition (Subtask 2), and diacritic restoration for spoken dialects (Subtask 3). A total of 44 teams registered, and during the testing phase, 100 valid submissions were received from eight unique teams. The distribution was as follows: 34 submissions for Subtask 1 "five teams{\ae}, 47 submissions for Subtask 2 "six teams", and 19 submissions for Subtask 3 "two teams". The best-performing systems achieved 79.8% accuracy on Subtask 1, 35.68/12.20 WER/CER (overall average) on Subtask 2, and 55/13 WER/CER on Subtask 3. These results highlight the ongoing challenges of Arabic dialect speech processing, particularly in dialect identification, recognition, and diacritic restoration. We also summarize the methods adopted by participating teams and briefly outline directions for future editions of NADI.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ahmed Amine Ben Abdallah, Ata Kabboudi, Amir Kanoun, and Salah Zaiem. 2023. https://arxiv.org/abs/2309.11327 Leveraging data collection and unsupervised learning for code-switched tunisian arabic automatic speech recognition . Preprint, arXiv:2309.11327
work page Pith review arXiv 2023
-
[4]
Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan, and Kareem Darwish. 2021. Qadi: Arabic dialect identification in the wild. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 1--10, Kyiv, Ukraine (Virtual). Association for Computational Linguistics
2021
-
[5]
Muhammad Abdul-Mageed, Hassan Alhuzali, and Mohamed Elaraby. 2018. You tweet what you speak: A city-level dataset of arabic dialects. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA)
2018
-
[6]
Muhammad Abdul-Mageed, AbdelRahim Elmadany, Chiyu Zhang, El Moatez Billah Nagoudi, Houda Bouamor, and Nizar Habash. 2023. Nadi 2023: The fourth nuanced arabic dialect identification shared task. arXiv preprint arXiv:2310.16117
work page Pith review arXiv 2023
-
[7]
Muhammad Abdul-Mageed, Amr Keleg, AbdelRahim Elmadany, Chiyu Zhang, Injy Hamed, Walid Magdy, Houda Bouamor, and Nizar Habash. 2024. Nadi 2024: The fifth nuanced arabic dialect identification shared task. arXiv preprint arXiv:2407.04910
work page Pith review arXiv 2024
-
[8]
Muhammad Abdul - Mageed, Chiyu Zhang, Houda Bouamor, and Nizar Habash. 2020. https://www.aclweb.org/anthology/2020.wanlp-1.9/ NADI 2020: The first nuanced arabic dialect identification shared task . In Proceedings of the Fifth Arabic Natural Language Processing Workshop, WANLP@COLING 2020, Barcelona, Spain (Online), December 12, 2020, pages 97--110. Assoc...
2020
Show all 92 references
-
[9]
Muhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim Elmadany, Houda Bouamor, and Nizar Habash. 2022. Nadi 2022: The third nuanced arabic dialect identification shared task. arXiv preprint arXiv:2210.09582
2022 arXiv
-
[10]
Muhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim Elmadany, and Lyle Ungar. 2020. Toward micro-dialect identification in diaglossic and code-switched environments. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5855--5876,...
2020
-
[11]
Elmadany, Houda Bouamor, and Nizar Habash
Muhammad Abdul - Mageed, Chiyu Zhang, AbdelRahim A. Elmadany, Houda Bouamor, and Nizar Habash. 2021. https://www.aclweb.org/anthology/2021.wanlp-1.28/ NADI 2021: The second nuanced arabic dialect identification shared task . In Proceedings of the Sixth Arabic Natural Language ...
2021
-
[12]
Badr Abdullah, Yusser Al-Ghussein, Zena Al-Khalili, Ömer Özyilmaz, Matias Valdenegro-Toro, Simon Ostermann, and Dietrich Klakow. 2025. Saarland-groningen at nadi 2025 shared task: Effective dialectal arabic speech processing under data constraints. In The Third Arabic Natural ...
2025
-
[13]
Kathrein Abu Kwaik, Motaz Saad, Stergios Chatzikyriakidis, and Simon Dobnik. 2018. Shami: A corpus of levantine arabic dialects. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)
2018
-
[14]
Maryam Khalifa Al Ali and Hanan Aldarmaki. 2024. https://aclanthology.org/2024.sigul-1.26 Mixat: A data set of bilingual emirati- E nglish speech . In Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages @ LREC-COLING 2024, pages 222...
2024
-
[15]
Abandah, Adham Alsharkawi, and Maha Dawas
Mohammad Al-Fetyani, Muhammad Al-Barham, Gheith A. Abandah, Adham Alsharkawi, and Maha Dawas. 2022. https://doi.org/10.1109/SLT54892.2023.10022652 MASC : Massive arabic speech corpus . In Proceedings of the 2022 IEEE Spoken Language Technology Workshop (SLT), page 1002206
2022
-
[16]
Rania Al-Sabbagh and Roxana Girju. 2012. Yadac: Yet another dialectal arabic corpus. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 2882--2889, Istanbul, Turkey. European Language Resources Association (ELRA)
2012
-
[17]
Nora Al-Twairesh, Rawan N. Al-Matham, Nora Madi, Nada Almugren, Al-Hanouf Al-Aljmi, Shahad Alshalan, Raghad Alshalan, Nafla Alrumayyan, Shams Al-Manea, Sumayah Bawazeer, Nourah Al-Mutlaq, Nada Almanea, Waad Bin Huwaymil, Dalal Alqusair, Reem Alotaibi, Suha Al-Senaydi, and Abee...
2018
-
[18]
Faris Alasmary, Orjuwan Zaafarani, and Ahmad Ghannam. 2024. https://arxiv.org/abs/2407.03236 Catt: Character-based arabic tashkeel transformer . Preprint, arXiv:2407.03236
2024 arXiv
-
[19]
Sanad ALBawwab and Omar Qawasmeh. 2025. Lahjati at nadi 2025 a ecapa-wavlm fusion with multi-stage optimization. In The Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou. Association for Computational Linguistics
2025
-
[20]
Hanan Aldarmaki and Ahmad Ghannam. 2023. Diacritic recognition performance in arabic asr. In Proc. Interspeech 2023, pages 361--365
2023
-
[21]
Ahmed Ali, Peter Bell, James Glass, Yacine Messaoui, Hamdy Mubarak, Steve Renals, and Yifan Zhang. 2016. https://doi.org/10.1109/SLT.2016.7846277 The mgb-2 challenge: Arabic multi-dialect broadcast media recognition . In 2016 IEEE Spoken Language Technology Workshop (SLT), pag...
2016
-
[22]
Ahmed Ali, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James Glass, Steve Renals, and Khalid Choukri. 2019. The mgb-5 challenge: Recognition and dialect identification of dialectal arabic speech. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop...
2019
-
[23]
Ahmed Ali, Stephan Vogel, and Steve Renals. 2017. Speech recognition challenge in the wild: Arabic mgb-3. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 316--322. IEEE
2017
-
[24]
Khalid Almeman, Mark Lee, and Ali Abdulrahman Almiman. 2013. https://doi.org/10.1109/ICCSPA.2013.6487288 Multi dialect arabic speech parallel corpora . In 2013 1st International Conference on Communications, Signal Processing, and their Applications (ICCSPA), pages 1--6
2013
-
[25]
Israa Alsarsour, Esraa Mohamed, Reem Suwaileh, and Tamer Elsayed. 2018. Dart: A large dataset of dialectal arabic tweets. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)
2018
-
[26]
Althobaiti
Maha J. Althobaiti. 2022. Creation of annotated country-level dialectal arabic resources: An unsupervised approach. Natural Language Engineering, 28(5):607--648
2022
-
[27]
Tyers, and Gregor Weber
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M. Tyers, and Gregor Weber. 2020. Common voice: A massively-multilingual speech corpus. In Proceedings of the 12th Conference on Language Resources...
2020
-
[28]
Rafiul Biswas, Shimaa Ibrahim, Mabrouka Bessghaier, Firoj Alam, and Wajdi Zaghouani
Kais Attia, Md. Rafiul Biswas, Shimaa Ibrahim, Mabrouka Bessghaier, Firoj Alam, and Wajdi Zaghouani. 2025. Marsadlab at nadi: Arabic dialect identification and speech recognition using ecapa-tdnn and whisper. In The Third Arabic Natural Language Processing Conference (ArabicNL...
2025
-
[29]
Matthew Baas, Benjamin van Niekerk , and Herman Kamper. 2023. https://doi.org/10.21437/Interspeech.2023-419 Voice conversion with just nearest neighbors . In Interspeech 2023, pages 2053--2057
2023 doi
-
[30]
As-Said Muh \'a mmad Badawi. 1973. Mustawayat al-arabiyya al-muasira fi Misr. Dar al-maarif
1973
-
[31]
Nurpeiis Baimukan, Houda Bouamor, and Nizar Habash. 2022. Hierarchical aggregation of dialectal data for arabic dialect identification. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (LREC 2022), pages 4586--4596, Marseille, France. European Lang...
2022
-
[32]
Lo \" c Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, and 1 others. 2023. Seamless: Multilingual expressive and streaming speech translation. arXiv preprint arXiv:2312.05187
2023 arXiv
-
[33]
Houda Bouamor, Nizar Habash, and Kemal Oflazer. 2014. A multidialectal parallel corpus of arabic. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 1240--1245, Reykjavik, Iceland. European Language Resources Association (ELRA)
2014
-
[34]
Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, and Kemal Oflazer. 2018. http://www.lrec-conf.org/proceedings/lrec2018/summaries/351.html The MADAR arabic dialect corpus...
2018
-
[35]
Kristen Brustad. 2000. The Syntax of Spoken Arabic: A Comparative Study of Moroccan, Egyptian, Syrian, and Kuwaiti Dialects. Georgetown University Press
2000
-
[36]
Shammur A Chowdhury, Ahmed Ali, Suwon Shon, and James R Glass. 2020. What does an end-to-end dialect identification model learn about non-dialectal information? In INTERSPEECH, pages 462--466
2020
-
[37]
Al-Natsheh, Houda Bouamor, Karim Bouzoubaa, Violetta Cavalli-Sforza, Samhaa R
Kareem Darwish, Nizar Habash, Mourad Abbas, Hend Al-Khalifa, Huseein T. Al-Natsheh, Houda Bouamor, Karim Bouzoubaa, Violetta Cavalli-Sforza, Samhaa R. El-Beltagy, Wassim El-Hajj, Mustafa Jarrar, and Hamdy Mubarak. 2021. https://doi.org/10.1145/3447735 A P anoramic survey of N ...
2021 doi
-
[38]
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020. https://doi.org/10.21437/Interspeech.2020-2650 Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification . In Interspeech 2020, pages 3830--3834
2020 doi
-
[39]
Mona Diab, Nizar Habash, Owen Rambow, Mohamed Altantawy, and Yassine Benajiba. 2010. Colaba: Arabic dialect annotation and processing. In Proceedings of the LREC Workshop on Semitic Language Processing, pages 66--74
2010
-
[40]
Amirbek Djanibekov, Hawau Olamide Toyin, Raghad Alshalan, Abdullah Alitr, and Hanan Aldarmaki. 2025. https://arxiv.org/abs/2411.05872 Dialectal coverage and generalization in arabic speech recognition . Preprint, arXiv:2411.05872
2025 arXiv
-
[41]
Mahmoud El-Haj. 2020. Habibi -- a multi dialect multi national arabic song lyrics corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 1318--1326, Marseille, France. European Language Resources Association (ELRA)
2020
-
[42]
Heba Elfardy, Mohamed Al-Badrashiny, and Mona Diab. 2014. Aida: Identifying code switching in informal arabic text. In Proceedings of the First Workshop on Computational Approaches to Code Switching, pages 94--101, Doha, Qatar. Association for Computational Linguistics
2014
-
[43]
Salman Elgamal, Ossama Obeid, Mhd Kabbani, Go Inoue, and Nizar Habash. 2024. https://doi.org/10.18653/v1/2024.acl-long.792 A rabic diacritics in the wild: Exploiting opportunities for improved diacritization . In Proceedings of the 62nd Annual Meeting of the Association for Co...
2024 doi
-
[44]
Haroun Elleuch, Salima Mdhaffar, Yannick Estève, and Fethi Bougares. 2025 . https://doi.org/ 10.21437/Interspeech.2025-884 ADI-20: Arabic Dialect Identification dataset and models . In Interspeech 2025 , pages 2775--2779
2025 doi
-
[45]
Haroun Elleuch, Youssef Saidi, Salima Mdhaffar, Yannick Estève, and Fethi Bougares. 2025. Elyadata & lia at nadi 2025: Asr and adi subtasks. In The Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou. Association for Computational Linguistics
2025
-
[46]
AbdelRahim Elmadany, Muhammad Abdul-Mageed, and 1 others. 2022. Arat5: Text-to-text transformers for arabic language generation. In Proceedings of the 60th annual meeting of the association for computational linguistics (Volume 1: Long papers), pages 628--647
2022
-
[47]
Abdelrahim Elmadany, Muhammad Abdul-Mageed, and 1 others. 2023 a . Octopus: A multitask model and toolkit for arabic natural language generation. In Proceedings of ArabicNLP 2023, pages 232--243
2023
-
[48]
AbdelRahim Elmadany, ElMoatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023 b . Orca: A challenging benchmark for arabic language understanding. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9559--9586, Toronto, Canada. Association for Computa...
2023
-
[49]
Mohamed Lotfy Elrefai. 2025. Unicorn at nadi 2025 subtask 3: Gemm3n-dr: Audio-text diacritic restoration via fine-tuned multimodal arabic llm. In The Third Arabic Natural Language Processing Conference (ArabicNLP 2025), Suzhou. Association for Computational Linguistics
2025
-
[50]
Ali Fadel, Ibraheem Tuffaha, Bara ' Al-Jawarneh, and Mahmoud Al-Ayyoub. 2019. https://doi.org/10.18653/v1/D19-5229 Neural A rabic text diacritization: State of the art results and a novel approach for machine translation . In Proceedings of the 6th Workshop on Asian Translatio...
2019 doi
-
[51]
Hassan Gadalla, Hanaa Kilany, Howaida Arram, Ashraf Yacoub, Alaa El-Habashi, Amr Shalaby, Krisjanis Karins, Everett Rowson, Robert MacIntyre, Paul Kingsbury, David Graff, and Cynthia McLemore. 1997. Callhome egyptian arabic transcripts ldc97t19. Web Download
1997
-
[52]
Ahmad Ghannam, Naif Alharthi, Faris Alasmary, Kholood Al Tabash, Shouq Sadah, and Lahouari Ghouti. 2025. Abjad ai at nadi 2025: Catt-whisper: Multimodal diacritic restoration using text and speech representations. In The Third Arabic Natural Language Processing Conference (Ara...
2025
-
[53]
Alex Graves, Santiago Fern \'a ndez, Faustino Gomez, and J \"u rgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning, pages 369--376
2006
-
[54]
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and 1 others. 2020. Conformer: Convolution-augmented transformer for speech recognition. arXiv preprint arXiv:2005.08100
2020 arXiv
-
[55]
Karim El Haff, Mustafa Jarrar, Tymaa Hammouda, and Fadi Zaraket. 2022. https://aclanthology.org/2022.lrec-1.82.pdf C urras + B aladi: T owards a L evantine C orpus . In Proceedings of the International Conference on Language Resources and Evaluation(LREC 2022), Marseille, France
2022
-
[56]
Nagham Hamad, Mohammed Khalilia, and Mustafa Jarrar. 2025. http://www.jarrar.info/publications/HKJ25.pdf K onooz: M ulti-domain M ulti-dialect C orpus for N amed E ntity R ecognition . In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, ...
2025
-
[57]
Injy Hamed, Ngoc Thang Vu, and Slim Abdennadher. 2020. https://aclanthology.org/2020.lrec-1.523/ A rz E n: A speech corpus for code-switched E gyptian A rabic- E nglish . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4237--4246, Marseille, F...
2020
-
[58]
Harrat, M
S. Harrat, M. Abbas, K. Meftouh, and K. Smaili. 2013. https://doi.org/10.21437/Interspeech.2013-373 Diacritics restoration for arabic dialect texts . In Interspeech 2013, pages 1429--1433
2013 doi
-
[59]
Salima Harrat, Karima Meftouh, Mourad Abbas, and Kamel Sma \"i li. 2014. Building resources for algerian arabic dialects. In INTERSPEECH 2014, 15th Annual Conference of the International Speech Communication Association, pages 2123--2127, Singapore. ISCA
2014
-
[60]
Richard S. Harrell. 1962. A Short Reference Grammar of Moroccan Arabic: With Audio CD. Georgetown Classics in Arabic Language and Linguistics. Georgetown University Press
1962
-
[61]
Clive Holes. 2004. Modern Arabic: Structures, Functions, and Varieties. Georgetown Classics in Arabic Language and Linguistics. Georgetown University Press
2004
-
[62]
Elsayed Issa, Mohammed AlShakhori, Reda AlBahrani, and Gus Hahn-Powell. 2021. Country-level arabic dialect identification using rnns with and without linguistic features. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 276--281, Kyiv, Ukraine (Vi...
2021
-
[63]
Mustafa Jarrar, Nizar Habash, Faeq Alrimawi, Diyam Akra, and Nasser Zalmout. 2016. Curras: An annotated corpus for the palestinian arabic dialect. Language Resources and Evaluation, pages 1--31
2016
-
[64]
Mustafa Jarrar, Fadi Zaraket, Tymaa Hammouda, Daanish Masood Alavi, and Martin Waahlisch. 2023. https://doi.org/1 10.1109/AICCSA59173.2023.10479250 L isan: Y emeni, I rqi, L ibyan, and S udanese A rabic D ialect C opora with M orphological A nnotations . In The 20th IEEE/ACS I...
2023
-
[65]
Salam Khalifa, Nizar Habash, Dana Abdulrahim, and Sara Hassan. 2016. A large scale corpus of gulf arabic. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4282--4289, Portorož, Slovenia. European Language Resources Asso...
2016
-
[66]
Taku Kudo and John Richardson. 2018. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226
2018 arXiv
-
[67]
Ajinkya Kulkarni, Atharva Kulkarni, Sara Abedalmon'em Mohammad Shatnawi, and Hanan Aldarmaki. 2023. https://doi.org/10.21437/Interspeech.2023-2224 Clartts: An open-source classical arabic text-to-speech corpus . In 2023 INTERSPEECH, pages 5511--5515
2023 doi
-
[68]
Mingfei Lau, Qian Chen, Yeming Fang, Tingting Xu, Tongzhou Chen, and Pavel Golik. 2025. https://doi.org/10.18653/v1/2025.acl-long.370 Data quality issues in multilingual speech datasets: The need for sociolinguistic awareness and proactive language planning . In Proceedings of...
2025 doi
-
[69]
Yooyoung Lee, Craig Greenberg, Lisa Mason, and Elliot Singer. 2022. Nist 2022 language recognition evaluation plan
2022
-
[70]
Shervin Malmasi, Marcos Zampieri, Nikola Ljube s i \'c , Preslav Nakov, Ahmed Ali, and J \"o rg Tiedemann. 2016. Discriminating between similar languages and A rabic dialect identification: A report on the third DSL shared task. In Proceedings of the third workshop on NLP for ...
2016
-
[71]
Karima Meftouh, Salima Harrat, Salma Jamoussi, Mourad Abbas, and Kamel Smaili. 2015. Machine translation experiments on padic: A parallel arabic dialect corpus. In Proceedings of the 29th Pacific Asia Conference on Language, Information and Computation, pages 26--34, Shanghai, China
2015
-
[72]
Hamdy Mubarak and Kareem Darwish. 2014. Using twitter to collect a multi-dialectal corpus of arabic. In Proceedings of the EMNLP 2014 Workshop on Arabic Natural Language Processing (ANLP), pages 1--7, Doha, Qatar. Association for Computational Linguistics
2014
-
[73]
El Moatez Billah Nagoudi, Ahmed El-Shangiti, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2023. Dolphin: A challenging and diverse benchmark for arabic nlg. arXiv preprint arXiv:2305.14989
2023 arXiv
-
[74]
Amal Nayouf, Mustafa Jarrar, Fadi zaraket, Tymaa Hammouda, and Mohamad-Bassam Kurdy. 2023. https://doi.org/10.18653/v1/2023.arabicnlp-1.2 N âbra: S yrian A rabic D ialects with M orphological A nnotations . In Proceedings of the 1st Arabic Natural Language Processing Conferenc...
2023 doi
-
[75]
Ossama Obeid, Mohammad Salameh, Houda Bouamor, and Nizar Habash. 2019. Adida: Automatic dialect identification for arabic. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 6--11, Minnea...
2019
- [76]
-
[77]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492--28518. PMLR
2023
-
[78]
Fatiha Sadat, Farzindar Kazemi, and Atefeh Farzindar. 2014. Automatic identification of arabic language varieties and dialects in social media. In Proceedings of the Second Workshop on Natural Language Processing for Social Media (SocialNLP), pages 22--27, Dublin, Ireland. Ass...
2014
-
[79]
Mohammad Salameh, Houda Bouamor, and Nizar Habash. 2018. Fine-grained arabic dialect identification. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1332--1344, Santa Fe, New Mexico, USA. Association for Computational Linguistics
2018
-
[80]
Mahmoud Salhab, Marwan Elghitany, Shameed Sait, Syed Sibghat Ullah, Mohammad Abusheikh, and Hasan Abusheikh. 2025 a . Advancing arabic speech recognition through large-scale weakly supervised learning. arXiv preprint arXiv:2504.12254
2025 arXiv
-
[81]
Mahmoud Salhab, Shameed Sait, Mohammad Abusheikh, and Hasan Abusheikh. 2025 b . Munsit at nadi 2025 shared task 2: Pushing the boundaries of multidialectal arabic asr with weakly supervised pretraining and continual supervised fine-tuning. In The Third Arabic Natural Language ...
2025
-
[82]
Sara Shatnawi, Sawsan Alqahtani, and Hanan Aldarmaki. 2024. Automatic restoration of diacritics for speech data sets. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Lo...
2024
-
[83]
Suwon Shon, Ahmed Ali, Younes Samih, Hamdy Mubarak, and James Glass. 2020. Adi17: A fine-grained arabic dialect identification dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8244--8248. IEEE
2020
-
[84]
Peter Sullivan, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2023. https://doi.org/10.21437/Interspeech.2023-1005 On the robustness of arabic speech dialect identification . In Interspeech 2023, pages 5326--5330
2023 doi
-
[85]
Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah, Chafei Mohamed Chafei, Ahmed Oumar El-Shangiti, Hiba Zayed, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, and 1 others. 2024. Casablanca: Data and models for multidialectal arabic speech recognition. arXiv ...
2024 arXiv
-
[86]
Hawau Toyin, Amirbek Djanibekov, Ajinkya Kulkarni, and Hanan Aldarmaki. 2023. Artst: Arabic text and speech transformer. In Proceedings of ArabicNLP 2023, pages 41--51
2023
-
[87]
Magdy, and Hanan Aldarmaki
Hawau Toyin, Rufael Marew, Humaid Alblooshi, Samar M. Magdy, and Hanan Aldarmaki. 2025 . https://doi.org/ 10.21437/Interspeech.2025-1550 ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis . In Interspeech 2025 , pages 4808--4812
2025 doi
-
[88]
o rgen Valk and Tanel Alum \
J \"o rgen Valk and Tanel Alum \"a e. 2021. Voxlingua107: a dataset for spoken language recognition. In 2021 IEEE Spoken Language Technology Workshop (SLT), pages 652--658. IEEE
2021
-
[89]
Abdul Waheed, Bashar Talafha, Peter Sullivan, Abdelrahim Elmadany, and Muhammad Abdul-Mageed. 2023. Voxarabica: A robust dialect-aware arabic speech recognition system. In Proceedings of ArabicNLP 2023, pages 441--449
2023
-
[90]
Wajdi Zaghouani and Anis Charfi. 2018. Arap-tweet: A large multi-dialect twitter corpus for gender, age and language variety identification. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Lang...
2018
-
[91]
Zaidan and Chris Callison-Burch
Omar F. Zaidan and Chris Callison-Burch. 2011. The arabic online commentary dataset: An annotated dataset of informal arabic with high dialectal content. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pa...
2011
-
[92]
Taha Zerrouki and Amar Balla. 2017. https://doi.org/10.1016/j.dib.2017.01.011 Tashkeela: Novel corpus of arabic vocalized texts, data for auto-diacritization systems . Data in Brief, 11
2017 doi
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.