Enhancing Arabic Dialectical Speech
Recognition with Large Language Models
Automatic
Speech Recognition (ASR) for Arabic faces unique challenges due to the
language’s diglossic nature, where Modern Standard Arabic (MSA) coexists with
diverse dialects that differ substantially at lexical, phonological, and
morphological levels. In this study, we present a comparative evaluation of
state-of-the-art ASR models, Whisper (medium and large) and Wav2Vec2, applied
to five Arabic varieties: Egyptian, Hijazi, Khaliji, Najdi, and MSA. We design
a two-phase experiment: in the first phase, 50 speech samples per dialect are
evaluated under three settings, raw transcriptions, and outputs corrected by
two large language models (GPT and ALLAM). In the second phase, the
best-performing ASR–LLM combination is scaled to the full test set for each
dialect. Performance is assessed using Word Error Rate (WER) and semantic
similarity distance. Results show that Whisper-Large combined with GPT
consistently outperforms other systems across dialects. Improvements are
especially pronounced in dialectal speech, while MSA achieves the lowest
baseline WER. These findings highlight both the persistent difficulty of
dialectal Arabic ASR and the effectiveness of LLM-based post-processing for
reducing recognition errors.
[1] A. Radford, J. W. Kim,
T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, "Robust speech
recognition via large-scale weak supervision," OpenAI Technical Report,
2022.
[2] A. Baevski, Y. Zhou, A.
Mohamed, and M. Auli, "wav2vec 2.0: A framework for self-supervised
learning of speech representations," in Advances in Neural Information
Processing Systems, vol. 33, 2020.
[3] Abdou and A. M. Moussa, "Arabic speech recognition: Challenges and
state of the art," in Computational Linguistics, Speech and Image
Processing for Arabic Language, 2019, pp. 1–27.
[4] A. Rahman, M. M. Kabir, M.
F. Mridha, M. Alatiyyah, H. F. Alhasson, and S. S. Alharbi, "Arabic speech
recognition: Advancement and challenges," IEEE Access, vol. 12,
pp. 39689–39716, 2024.
[5] M. Afify, R. Sarikaya,
H.-K. J. Kuo, L. Besacier, and Y. Gao, "On the use of morphological
analysis for dialectal Arabic speech recognition," in Proc.
Interspeech 2006, 2006, paper 1444-Mon2A2O.2, doi:
10.21437/Interspeech.2006-87.
[6] S. Alharbi et al.,
"SADA: Saudi audio dataset for Arabic," in Proc. IEEE Int. Conf.
Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10286–10290.
[7] A. Ali et al.,
"The MGB-2 challenge: Arabic multi-dialect broadcast media
recognition," in Proc. IEEE Spoken Language Technology Workshop (SLT),
2016, pp. 279–284, doi: 10.1109/SLT.2016.7846277.
[8] D. Wang, X. Wang, and
S. Lv, "An overview of end-to-end automatic speech recognition," Symmetry,
vol. 11, no. 8, p. 1018, 2019.
[9] N. Zerari, S.
Abdelhamid, H. Bouzgou, and C. Raymond, "Bidirectional deep architecture
for Arabic speech recognition," Open Computer Science, vol. 9,
pp. 92–102, 2019.
[10] A. Graves, S.
Fern?ndez, F. Gomez, and J. Schmidhuber, "Connectionist temporal
classification: Labelling unsegmented sequence data with recurrent neural
networks," in Proc. 23rd Int. Conf. Machine Learning (ICML),
2006, pp. 369–376.
[11] W. Chan, N. Jaitly, Q. V.
Le, and O. Vinyals, "Listen, attend and spell: A neural network for large
vocabulary conversational speech recognition," in Proc. IEEE Int.
Conf. Acoustics, Speech and Signal Processing (ICASSP), 2016, pp.
4960–4964.
[12] A. Waheed, B. Talafha,
P. Sullivan, A. Elmadany, and M. Abdul-Mageed, "VoxArabica: A robust
dialect-aware Arabic speech recognition system," in Proc. ArabicNLP
2023 Workshop, 2023, pp. 1–14.
[13] B. Talafha, A. Waheed,
and M. Abdul-Mageed, "N-shot benchmarking of Whisper on diverse Arabic
speech recognition," arXiv preprint, 2023.
[14] R. Sachdev, Z. Q.
Wang, and C. H. H. Yang, "Evolutionary prompt design for LLM-based
post-ASR error correction," arXiv preprint arXiv:2407.16370, 2024.
[15] R. Ma, M. Qian, M.
Gales, and K. Knill, "ASR error correction using large language
models," IEEE Trans. Audio, Speech, Language Process., 2025.
[16] R. Ma, M. Qian, P.
Manakul, M. Gales, and K. Knill, "Can generative large language models
perform ASR error correction?" arXiv preprint arXiv:2307.04172, 2023.
[17] C. H. H. Yang, Y. Gu,
Y. C. Liu, S. Ghosh, I. Bulyko, and A. Stolcke, "Generative speech
recognition error correction with large language models and task-activating
prompting," in Proc. IEEE Automatic Speech Recognition and
Understanding Workshop (ASRU), 2023, pp. 1–8.
[18] Z. Min and J. Wang,
"Exploring the integration of large language models into automatic speech
recognition systems: An empirical study," in Proc. Int. Conf. Neural
Information Processing (ICONIP), Singapore, 2023, pp. 69–84.
[19] S. Li, C. Chen, C. Y.
Kwok, C. Chu, E. S. Chng, and H. Kawai, "Investigating ASR error
correction with large language model and multilingual 1-best hypotheses,"
in Proc. Interspeech, 2024, pp. 1315–1319.
[20] N. M. Matasyoh, R. A.
Zeineldin, and F. Mathis-Ullrich, "Optimising speech recognition using
LLMs: An application in the surgical domain," Current Directions in
Biomedical Engineering, vol. 10, no. 1, pp. 45–48, Sep. 2024.
[21] S. Alharbi, R.
Binmuqbil, A. Ali, R. Aloraini, S. Bari, A. Alowisheq, and Y. Alonaizan,
"Leveraging LLM for augmenting textual data in code-switching ASR: Arabic
as an example," in Proc. SynData4GenAI, 2024.
[22] T. Inoue, M. Elmahdy,
A. Ali, and P. Bell, "Comparative study of end-to-end and hybrid ASR
systems for Arabic broadcast speech," in Proc. IEEE Automatic Speech
Recognition and Understanding Workshop (ASRU), 2021, pp. 1085–1092.
[23] Y. Zhang, X. Wang, J.
Xu, and H. Wang, "A comparative study of deep learning models for Arabic
speech recognition," IEEE Access, vol. 10, pp. 10945–10953, 2022.
[24] S. Ding, A. Guo, A.-V.
Nguyen, A. Vidon, J. Wang, and M. Zhou, "Data science capstone project:
Robustness of Whisper and Wav2Vec2.0," 2023.
[25] Y. Wang, A. Alhmoud,
and M. Alqurishi, "Open universal Arabic ASR leaderboard," arXiv
preprint, 2024.
[26] A. C. Morris, V.
Maier, and P. Green, "From WER and RIL to MER and WIL: Improved evaluation
measures for connected speech recognition," in Proc. Interspeech,
2004.
[27] A. Sahyoun and S.
Shehata, "AraDiaWER: An explainable metric for dialectical Arabic
ASR," in Proc. 2nd Workshop on NLP Applications to Field Linguistics,
Dubrovnik, Croatia, 2023, pp. 64–73.
[28] N. Reimers and I.
Gurevych, "Sentence-BERT: Sentence embeddings using Siamese
BERT-networks," in Proc. Conf. Empirical Methods in Natural Language
Processing (EMNLP/IJCNLP), 2019.
[29] M. Alian and A.
Awajan, "Arabic semantic similarity approaches - review," in Proc.
Int. Arab Conf. Information Technology (ACIT), 2018, pp. 1–6, doi:
10.1109/ACIT.2018.8672665.
[30] T. Brown et al.,
"Language models are few-shot learners," in Advances in Neural
Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.
F. Balcan, and H. Lin, Eds., vol. 33, 2020, pp. 1877–1901.
[31] M. S. Bari et al.,
"ALLaM: Large language models for Arabic and English," arXiv
preprint, 2024.
[32] K. Almeman and M. Lee,
"A comparison of Arabic speech recognition for multi-dialect vs. specific
dialects," in Proc. 7th Int. Conf. Speech Technology and
Human-Computer Dialogue (SpeD), Cluj-Napoca, Romania, 2013.
[33] M. Naderi et al.,
"Towards interfacing large language models with ASR systems using
confidence measures and prompting," arXiv preprint arXiv:2407.19549, 2024.
[34] ?. T. ?zyilmaz, M.
Coler, and M. Valdenegro-Toro, "Overcoming data scarcity in
multi-dialectal Arabic ASR via Whisper fine-tuning," arXiv preprint
arXiv:2506.02627, 2025.