معالجة العربية ولهجاتها
مجموعات بيانات ونماذج تفهم العربية كما تُكتب فعلًا على الإنترنت، بلهجاتها وطرائق كتابتها وتناوبها بين اللغات.
- تمييز اللهجات
- لغات قليلة الموارد
- العربيزي
البحث
أدرس كيف تتعامل تقنيات اللغة مع العربية بكل تنوّعها، وكيف تنتقل المعلومات — والمعلومات المضلّلة — عبر الإعلام الناطق بالعربية. تجد أدناه مجالات عملي الأساسية، والمشاريع التي أقودها حاليًا، والقائمة الكاملة لمنشوراتي.
مجموعات بيانات ونماذج تفهم العربية كما تُكتب فعلًا على الإنترنت، بلهجاتها وطرائق كتابتها وتناوبها بين اللغات.
كيف تنتشر الادعاءات الزائفة عبر المنصات الناطقة بالعربية، ولماذا تُخفق أدوات الكشف المدرّبة على الإنجليزية في رصدها.
تدقيق تقنيات اللغة بحثًا عن التحيّز، والمساهمة في سياسات لا تجعل العالم الناطق بالعربية أمرًا ثانويًا.
مناهج حاسوبية لقراءة الصحف العربية التاريخية وحفظ السجل الرقمي للمنطقة.
أعمال ممولة قيد التنفيذ، بالتعاون مع باحثين من أربع مؤسسات.
بناء مدوّنة من 40 مليون كلمة من العربية المحكية في 18 دولة، مجموعة بموافقة أصحابها، مع أدوات للباحثين والشركات الناشئة.
تتبّع انتقال المعلومات الصحية المضلّلة عبر تطبيقات المراسلة العربية خلال الأزمات الصحية، واختبار تدخلات بالتعاون مع مدقّقي الأخبار.
التعرّف الضوئي والبحث في 120 ألف صفحة من صحف بلاد الشام (1908–1948)، متاحة الآن مجانًا للمؤرخين.
ابحث بالعنوان أو اسم المؤلف المشارك أو جهة النشر، وانسخ مرجع BibTeX بنقرة واحدة.
عرض 10 من 10
2026
L. Haddad, O. Nasser, J. Kim
Journal of Computational Media Studies, 12(2)
Misinformation classifiers for Arabic are overwhelmingly trained on Modern Standard Arabic, yet most viral claims circulate in dialect. We introduce a benchmark of 38,000 fact-checked claims across nine dialect groups and show that dialect-aware pre-training reduces false negatives by 31% relative to strong multilingual baselines.
2025
L. Haddad, S. Benali, R. Aziz, M. Farouk
Proceedings of the Arabic NLP Conference (ArabicNLP 2025)
We describe the design and first release of LAHJA, a 12-million-word corpus of dialectal Arabic contributed by 4,100 volunteers with explicit, revocable consent. We detail our collection protocol, annotation scheme, and baseline results for dialect identification and sentiment analysis.
2025
ل. حدّاد، ع. الخطيب
مجلة اللسانيات العربية الحاسوبية، المجلد 7، العدد 1
تقيس هذه الدراسة أداء خمسة نماذج لغوية كبيرة على مهام الفهم والتوليد بخمس لهجات عربية، وتُظهر فجوة في الأداء تصل إلى 24 نقطة بين الفصحى واللهجات المغاربية. نقترح إطارًا للتقييم العادل بين اللهجات ونناقش أثره على الخدمات الرقمية الحكومية.
2024
L. Haddad, T. Ofori, D. Mansour
Proceedings of the ACM Conference on Fairness in Computing (FairComp '24)
Using 60,000 paired posts, we measure how automated moderation systems treat equivalent content written in English, Modern Standard Arabic, and Levantine Arabic. Dialectal posts were removed at 2.3× the rate of their English equivalents, with disproportionate impact on news and advocacy accounts.
2024
L. Haddad, N. Sayegh
Digital Scholarship in the Humanities, 39(4)
We report on a three-year collaboration to digitise 120,000 pages of Levantine newspapers (1908–1948). Beyond character error rates, we examine how historians actually search and cite OCR'd sources, and propose design principles for archives serving Arabic-script collections.
2023
L. Haddad
In A. Rahman & P. Cole (Eds.), The Handbook of Language and Digital Society. Meridian Academic Press
This chapter traces how decisions about which variety of Arabic to support — in keyboards, spell-checkers, and machine translation — encode older debates about authenticity, nationhood, and modernity, and argues for a pluralist approach to language technology design.
2023
ل. حدّاد، ي. العمري، هـ. سليمان
المجلة العربية لبحوث الإعلام والاتصال، العدد 42
تحلل الدراسة 1.2 مليون رسالة في 340 مجموعة مراسلة عامة ناطقة بالعربية، وتحدد أنماط انتشار الادعاءات الصحية الزائفة ودور «العُقد الجسرية» في نقلها بين المجتمعات. تقدم النتائج توصيات عملية لمنصات التحقق.
2022
R. Aziz, L. Haddad
Findings of the Conference on Empirical Methods in Language Processing
Arabizi — Arabic written in Latin characters and numerals — remains widespread yet poorly supported. We release a parallel Arabizi–Arabic corpus and a character-level transliteration model that improves downstream sentiment classification by 14 F1 points.
2021
L. Haddad
Northbridge University Press
Drawing on a decade of data from forums, blogs, and social media, this book shows how Arabic speakers use spelling, script, and dialect online to signal identity and belonging — and what this means for the technologies that attempt to read them.
2019
L. Haddad, K. Lindqvist
Transactions on Language Resources, 7
We present a hierarchical approach to identifying Arabic dialects at the city level, combining regional and local classifiers. The model reaches 71% accuracy across 25 cities and remains a widely used baseline.