Arabic NLP & Dialects
Datasets and models that understand Arabic as it's actually written online — across dialects, scripts, and code-switching.
- Dialect identification
- Low-resource NLP
- Arabizi
Currently: Visiting Fellow, Centre for Digital Society (2026–27)
Associate Professor of Computational LinguisticsMashriq University, Amman
I study how Arabic-language media and AI systems shape what people come to believe — and I write about it for readers beyond the academy.
Research
My work sits between computer science and the humanities: building tools for Arabic, and asking who they serve.
Datasets and models that understand Arabic as it's actually written online — across dialects, scripts, and code-switching.
How false claims spread across Arabic-language platforms, and why detection tools trained on English miss them.
Auditing language technologies for bias, and shaping policy so the Arabic-speaking world isn't an afterthought.
Computational methods for reading historical Arabic newspapers and preserving the region's born-digital record.
Writing
Long-form pieces for readers beyond the academy.
More than 400 million people speak Arabic, but the Arabic most AI systems learn is one almost nobody speaks at home. That gap has consequences.
9 min read

We traced a single false health claim across 340 group chats. Its journey reveals why fact-checks so often arrive too late.
7 min read
Entire chapters of the region's recent history lived on blogs and forums that no longer exist. Deciding what to save is a political act.
8 min read
“AI outperforms doctors.” “Model understands language.” Five questions to ask before you share the next breakthrough.
6 min read
Publications
Recent peer-reviewed work. The full list is searchable on the research page.
2026
L. Haddad, O. Nasser, J. Kim
Journal of Computational Media Studies, 12(2)
Misinformation classifiers for Arabic are overwhelmingly trained on Modern Standard Arabic, yet most viral claims circulate in dialect. We introduce a benchmark of 38,000 fact-checked claims across nine dialect groups and show that dialect-aware pre-training reduces false negatives by 31% relative to strong multilingual baselines.
2025
L. Haddad, S. Benali, R. Aziz, M. Farouk
Proceedings of the Arabic NLP Conference (ArabicNLP 2025)
We describe the design and first release of LAHJA, a 12-million-word corpus of dialectal Arabic contributed by 4,100 volunteers with explicit, revocable consent. We detail our collection protocol, annotation scheme, and baseline results for dialect identification and sentiment analysis.
2025
ل. حدّاد، ع. الخطيب
مجلة اللسانيات العربية الحاسوبية، المجلد 7، العدد 1
تقيس هذه الدراسة أداء خمسة نماذج لغوية كبيرة على مهام الفهم والتوليد بخمس لهجات عربية، وتُظهر فجوة في الأداء تصل إلى 24 نقطة بين الفصحى واللهجات المغاربية. نقترح إطارًا للتقييم العادل بين اللهجات ونناقش أثره على الخدمات الرقمية الحكومية.
2024
L. Haddad, T. Ofori, D. Mansour
Proceedings of the ACM Conference on Fairness in Computing (FairComp '24)
Using 60,000 paired posts, we measure how automated moderation systems treat equivalent content written in English, Modern Standard Arabic, and Levantine Arabic. Dialectal posts were removed at 2.3× the rate of their English equivalents, with disproportionate impact on news and advocacy accounts.
Essays and commentary have appeared in
Newsletter
One carefully written essay a month, plus notes on new research worth your time. Unsubscribe anytime.