Arabic NLP & Dialects
Datasets and models that understand Arabic as it's actually written online — across dialects, scripts, and code-switching.
- Dialect identification
- Low-resource NLP
- Arabizi
Research
I study how language technologies handle Arabic in all its variety, and how information — and misinformation — moves through Arabic-language media. Below are my core areas, the projects I'm leading now, and a complete list of publications.
Datasets and models that understand Arabic as it's actually written online — across dialects, scripts, and code-switching.
How false claims spread across Arabic-language platforms, and why detection tools trained on English miss them.
Auditing language technologies for bias, and shaping policy so the Arabic-speaking world isn't an afterthought.
Computational methods for reading historical Arabic newspapers and preserving the region's born-digital record.
Funded work in progress, with collaborators across four institutions.
Building a 40-million-word, consent-based corpus of dialectal Arabic from 18 countries, with tools for researchers and startups.
Tracing how health misinformation moved through Arabic messaging apps during public-health crises, and testing interventions with fact-checkers.
OCR and search for 120,000 pages of Levantine newspapers (1908–1948), now freely available to historians.
Search by title, co-author, or venue. Copy a BibTeX entry with one click.
Showing 10 of 10
2026
L. Haddad, O. Nasser, J. Kim
Journal of Computational Media Studies, 12(2)
Misinformation classifiers for Arabic are overwhelmingly trained on Modern Standard Arabic, yet most viral claims circulate in dialect. We introduce a benchmark of 38,000 fact-checked claims across nine dialect groups and show that dialect-aware pre-training reduces false negatives by 31% relative to strong multilingual baselines.
2025
L. Haddad, S. Benali, R. Aziz, M. Farouk
Proceedings of the Arabic NLP Conference (ArabicNLP 2025)
We describe the design and first release of LAHJA, a 12-million-word corpus of dialectal Arabic contributed by 4,100 volunteers with explicit, revocable consent. We detail our collection protocol, annotation scheme, and baseline results for dialect identification and sentiment analysis.
2025
ل. حدّاد، ع. الخطيب
مجلة اللسانيات العربية الحاسوبية، المجلد 7، العدد 1
تقيس هذه الدراسة أداء خمسة نماذج لغوية كبيرة على مهام الفهم والتوليد بخمس لهجات عربية، وتُظهر فجوة في الأداء تصل إلى 24 نقطة بين الفصحى واللهجات المغاربية. نقترح إطارًا للتقييم العادل بين اللهجات ونناقش أثره على الخدمات الرقمية الحكومية.
2024
L. Haddad, T. Ofori, D. Mansour
Proceedings of the ACM Conference on Fairness in Computing (FairComp '24)
Using 60,000 paired posts, we measure how automated moderation systems treat equivalent content written in English, Modern Standard Arabic, and Levantine Arabic. Dialectal posts were removed at 2.3× the rate of their English equivalents, with disproportionate impact on news and advocacy accounts.
2024
L. Haddad, N. Sayegh
Digital Scholarship in the Humanities, 39(4)
We report on a three-year collaboration to digitise 120,000 pages of Levantine newspapers (1908–1948). Beyond character error rates, we examine how historians actually search and cite OCR'd sources, and propose design principles for archives serving Arabic-script collections.
2023
L. Haddad
In A. Rahman & P. Cole (Eds.), The Handbook of Language and Digital Society. Meridian Academic Press
This chapter traces how decisions about which variety of Arabic to support — in keyboards, spell-checkers, and machine translation — encode older debates about authenticity, nationhood, and modernity, and argues for a pluralist approach to language technology design.
2023
ل. حدّاد، ي. العمري، هـ. سليمان
المجلة العربية لبحوث الإعلام والاتصال، العدد 42
تحلل الدراسة 1.2 مليون رسالة في 340 مجموعة مراسلة عامة ناطقة بالعربية، وتحدد أنماط انتشار الادعاءات الصحية الزائفة ودور «العُقد الجسرية» في نقلها بين المجتمعات. تقدم النتائج توصيات عملية لمنصات التحقق.
2022
R. Aziz, L. Haddad
Findings of the Conference on Empirical Methods in Language Processing
Arabizi — Arabic written in Latin characters and numerals — remains widespread yet poorly supported. We release a parallel Arabizi–Arabic corpus and a character-level transliteration model that improves downstream sentiment classification by 14 F1 points.
2021
L. Haddad
Northbridge University Press
Drawing on a decade of data from forums, blogs, and social media, this book shows how Arabic speakers use spelling, script, and dialect online to signal identity and belonging — and what this means for the technologies that attempt to read them.
2019
L. Haddad, K. Lindqvist
Transactions on Language Resources, 7
We present a hierarchical approach to identifying Arabic dialects at the city level, combining regional and local classifiers. The model reaches 71% accuracy across 25 cities and remains a widely used baseline.