Islamic University Journal of Applied Sciences

Enhancing Arabic Dialectical Speech Recognition with Large Language Models

Rehab Albeladi

Keywords: Automatic Speech Recognition; Large Language Models; Arabic Dialects.

Major: Engineering

Sub Major: Communication theory and applications

https://doi.org/10.63070/jesc.2026.028; Received 15 February 2026; Revised 20 April 2026; Accepted 24 April 2026; Available online 30 April 2026.
DownloadPDF
Abstract

Automatic Speech Recognition (ASR) for Arabic faces unique challenges due to the language’s diglossic nature, where Modern Standard Arabic (MSA) coexists with diverse dialects that differ substantially at lexical, phonological, and morphological levels. In this study, we present a comparative evaluation of state-of-the-art ASR models, Whisper (medium and large) and Wav2Vec2, applied to five Arabic varieties: Egyptian, Hijazi, Khaliji, Najdi, and MSA. We design a two-phase experiment: in the first phase, 50 speech samples per dialect are evaluated under three settings, raw transcriptions, and outputs corrected by two large language models (GPT and ALLAM). In the second phase, the best-performing ASR–LLM combination is scaled to the full test set for each dialect. Performance is assessed using Word Error Rate (WER) and semantic similarity distance. Results show that Whisper-Large combined with GPT consistently outperforms other systems across dialects. Improvements are especially pronounced in dialectal speech, while MSA achieves the lowest baseline WER. These findings highlight both the persistent difficulty of dialectal Arabic ASR and the effectiveness of LLM-based post-processing for reducing recognition errors.

References

[1] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, "Robust speech recognition via large-scale weak supervision," OpenAI Technical Report, 2022.

[2] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, "wav2vec 2.0: A framework for self-supervised learning of speech representations," in Advances in Neural Information Processing Systems, vol. 33, 2020.

[3] Abdou and A. M. Moussa, "Arabic speech recognition: Challenges and state of the art," in Computational Linguistics, Speech and Image Processing for Arabic Language, 2019, pp. 1–27.

[4] A. Rahman, M. M. Kabir, M. F. Mridha, M. Alatiyyah, H. F. Alhasson, and S. S. Alharbi, "Arabic speech recognition: Advancement and challenges," IEEE Access, vol. 12, pp. 39689–39716, 2024.

[5] M. Afify, R. Sarikaya, H.-K. J. Kuo, L. Besacier, and Y. Gao, "On the use of morphological analysis for dialectal Arabic speech recognition," in Proc. Interspeech 2006, 2006, paper 1444-Mon2A2O.2, doi: 10.21437/Interspeech.2006-87.

[6] S. Alharbi et al., "SADA: Saudi audio dataset for Arabic," in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10286–10290.

[7] A. Ali et al., "The MGB-2 challenge: Arabic multi-dialect broadcast media recognition," in Proc. IEEE Spoken Language Technology Workshop (SLT), 2016, pp. 279–284, doi: 10.1109/SLT.2016.7846277.

[8] D. Wang, X. Wang, and S. Lv, "An overview of end-to-end automatic speech recognition," Symmetry, vol. 11, no. 8, p. 1018, 2019.

[9] N. Zerari, S. Abdelhamid, H. Bouzgou, and C. Raymond, "Bidirectional deep architecture for Arabic speech recognition," Open Computer Science, vol. 9, pp. 92–102, 2019.

[10] A. Graves, S. Fern?ndez, F. Gomez, and J. Schmidhuber, "Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks," in Proc. 23rd Int. Conf. Machine Learning (ICML), 2006, pp. 369–376.

[11] W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, "Listen, attend and spell: A neural network for large vocabulary conversational speech recognition," in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 4960–4964.

[12] A. Waheed, B. Talafha, P. Sullivan, A. Elmadany, and M. Abdul-Mageed, "VoxArabica: A robust dialect-aware Arabic speech recognition system," in Proc. ArabicNLP 2023 Workshop, 2023, pp. 1–14.

[13] B. Talafha, A. Waheed, and M. Abdul-Mageed, "N-shot benchmarking of Whisper on diverse Arabic speech recognition," arXiv preprint, 2023.

[14] R. Sachdev, Z. Q. Wang, and C. H. H. Yang, "Evolutionary prompt design for LLM-based post-ASR error correction," arXiv preprint arXiv:2407.16370, 2024.

[15] R. Ma, M. Qian, M. Gales, and K. Knill, "ASR error correction using large language models," IEEE Trans. Audio, Speech, Language Process., 2025.

[16] R. Ma, M. Qian, P. Manakul, M. Gales, and K. Knill, "Can generative large language models perform ASR error correction?" arXiv preprint arXiv:2307.04172, 2023.

[17] C. H. H. Yang, Y. Gu, Y. C. Liu, S. Ghosh, I. Bulyko, and A. Stolcke, "Generative speech recognition error correction with large language models and task-activating prompting," in Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023, pp. 1–8.

[18] Z. Min and J. Wang, "Exploring the integration of large language models into automatic speech recognition systems: An empirical study," in Proc. Int. Conf. Neural Information Processing (ICONIP), Singapore, 2023, pp. 69–84.

[19] S. Li, C. Chen, C. Y. Kwok, C. Chu, E. S. Chng, and H. Kawai, "Investigating ASR error correction with large language model and multilingual 1-best hypotheses," in Proc. Interspeech, 2024, pp. 1315–1319.

[20] N. M. Matasyoh, R. A. Zeineldin, and F. Mathis-Ullrich, "Optimising speech recognition using LLMs: An application in the surgical domain," Current Directions in Biomedical Engineering, vol. 10, no. 1, pp. 45–48, Sep. 2024.

[21] S. Alharbi, R. Binmuqbil, A. Ali, R. Aloraini, S. Bari, A. Alowisheq, and Y. Alonaizan, "Leveraging LLM for augmenting textual data in code-switching ASR: Arabic as an example," in Proc. SynData4GenAI, 2024.

[22] T. Inoue, M. Elmahdy, A. Ali, and P. Bell, "Comparative study of end-to-end and hybrid ASR systems for Arabic broadcast speech," in Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2021, pp. 1085–1092.

[23] Y. Zhang, X. Wang, J. Xu, and H. Wang, "A comparative study of deep learning models for Arabic speech recognition," IEEE Access, vol. 10, pp. 10945–10953, 2022.

[24] S. Ding, A. Guo, A.-V. Nguyen, A. Vidon, J. Wang, and M. Zhou, "Data science capstone project: Robustness of Whisper and Wav2Vec2.0," 2023.

[25] Y. Wang, A. Alhmoud, and M. Alqurishi, "Open universal Arabic ASR leaderboard," arXiv preprint, 2024.

[26] A. C. Morris, V. Maier, and P. Green, "From WER and RIL to MER and WIL: Improved evaluation measures for connected speech recognition," in Proc. Interspeech, 2004.

[27] A. Sahyoun and S. Shehata, "AraDiaWER: An explainable metric for dialectical Arabic ASR," in Proc. 2nd Workshop on NLP Applications to Field Linguistics, Dubrovnik, Croatia, 2023, pp. 64–73.

[28] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence embeddings using Siamese BERT-networks," in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP/IJCNLP), 2019.

[29] M. Alian and A. Awajan, "Arabic semantic similarity approaches - review," in Proc. Int. Arab Conf. Information Technology (ACIT), 2018, pp. 1–6, doi: 10.1109/ACIT.2018.8672665.

[30] T. Brown et al., "Language models are few-shot learners," in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33, 2020, pp. 1877–1901.

[31] M. S. Bari et al., "ALLaM: Large language models for Arabic and English," arXiv preprint, 2024.

[32] K. Almeman and M. Lee, "A comparison of Arabic speech recognition for multi-dialect vs. specific dialects," in Proc. 7th Int. Conf. Speech Technology and Human-Computer Dialogue (SpeD), Cluj-Napoca, Romania, 2013.

[33] M. Naderi et al., "Towards interfacing large language models with ASR systems using confidence measures and prompting," arXiv preprint arXiv:2407.19549, 2024.

[34] ?. T. ?zyilmaz, M. Coler, and M. Valdenegro-Toro, "Overcoming data scarcity in multi-dialectal Arabic ASR via Whisper fine-tuning," arXiv preprint arXiv:2506.02627, 2025.