Skip to content
Ahmad
All articles
AI & Language

Why Arabic Dialects Still Break Most AI Models

More than 400 million people speak Arabic, but the Arabic most AI systems learn is one almost nobody speaks at home. That gap has consequences.

Ahmad Abu Hasan

9 min read

Originally published in Meridian Magazine

Ask a popular chatbot a question in Egyptian Arabic and it will usually answer — in formal Modern Standard Arabic, the language of news bulletins and official letters. It understood you, more or less. But it didn't speak your language back. For casual use, that's a quirk. For health information, legal advice, or content moderation, it's a real problem.

A language with two registers

Linguists call it diglossia: a formal written standard sits alongside a range of spoken dialects that differ in vocabulary, grammar, and pronunciation. A Moroccan and an Iraqi can both read the same newspaper, but their everyday speech can be as different as Spanish and Italian.

For decades, that split didn't matter much to computers, because the written web was mostly formal. Social media changed that. Today, most Arabic text produced every day is dialectal, often mixed with English or French, and frequently written in Latin letters.

The data problem

Language models learn from what they read. The largest Arabic datasets are dominated by news sites, Wikipedia, and religious texts — all written in the standard register. Dialects appear, but thinly, and some are almost absent. In our recent study, model accuracy on Maghrebi dialects trailed the standard by as much as 24 points.

When a system only understands the formal register, it ends up misreading the people who most need to be heard.

  • Content moderation removes harmless dialectal posts more often than equivalent English ones.
  • Voice assistants fail more often for older speakers and rural communities.
  • Misinformation written in dialect slips past detection tools trained on formal text.

What would actually help

The fix isn't a single bigger model. It's better, fairer data — collected with consent, balanced across regions, and documented well enough that others can audit it. It also means evaluating systems on the Arabic people actually use, not just on exam-style benchmarks.

None of this is exotic. It's the same care we'd expect for any product serving hundreds of millions of people. The question is whether the companies building these tools will treat Arabic speakers as a core audience or as an edge case.

Share