Abdulmuizz Khalak, Abderrahmane Issam, Gerasimos Spanakis
10 min
Abstract
Arabic Language Models (LMs) are pretrained predominately on Modern Standard Arabic (MSA) and are expected to transfer to its dialects. While MSA as the standard written variety is commonly used in formal settings, people speak and write online in various dialects that are spread across the Arab region. This poses limitations for Arabic LMs, since its dialects vary in their similarity to MSA. In this work we study cross-lingual transfer of Arabic models using probing on 3 Natural Language Processing (NLP) Tasks, and representational similarity. Our results indicate that transfer is possible but disproportionate across dialects, which we find to be partially explained by their geographic proximity. Furthermore, we find evidence for negative interference in models trained to support all Arabic dialects. This questions their degree of similarity, and raises concerns for cross-lingual transfer in Arabic models.
Alex: That covers grammar, spotting key info, and emotions—solid checks for real-world use. But how do they link this to geography?
Sam: They picked Yemen as a stand-in spot for MSA's roots, based on studies showing its dialect keeps old Arabic traits close to the formal version. Then, they measured straight-line distances from Yemen to each dialect's region on a map. The closer the spot, the better the probing scores dropped off predictably, suggesting geography shapes how much knowledge carries over.
Alex: So distance acts like a fade on the signal—makes sense for accents blending gradually.
Sam: Exactly. This proxy lets them test if AI mirrors the dialect continuum idea, where nearby varieties share more traits. It points to a clear pattern: transfer weakens with miles, even within one language family.
Alex: And they backed this with actual map math?
Sam: They did—straight-line distances from Yemen to dialect regions correlated with probing scores. But first, to ensure clean data, they filtered probing datasets from noisy online sources. Since those often mix dialects or sneak in formal Arabic, the team ran automatic checks with a tool that guesses the dialect's origin, like city or country, and kept only matching sentences. This sharpened the inputs for fair tests.
Alex: So like double-checking ingredients before cooking to avoid wrong flavors mixing in.
Sam: Exactly. For deeper checks on how models "think" alike, they used parallel sentences—same ideas written in formal Arabic and 25 city dialects—from a collection called the MADAR corpus. They fed these to different AI models, all of the same BERT-style family to keep comparisons even—BERT-style meaning those transformer brains that process whole sentences at once.
Alex: And what kinds did they pick?
Sam: They compared one trained purely on formal Arabic, one on mixed dialects, and several tuned just to single dialects—like Saudi, Egyptian, or North African ones such as Algerian or Moroccan. Probing showed the formal model tops charts on formal text for grammar, names, and feelings. Dialect-tuned ones shine on their home turf, but only if trained on big data piles; tiny ones lag. Saudi and Egyptian models came closest to formal performance on formal text, fitting their nearer geography. North African ones struggled most, showing bigger gaps in word forms and structure.
Alex: Huh, so scale of training data is key—like needing lots of practice reps in sports to compete broadly.
Sam: Right. Overall, this suggests distance predicts not just quiz scores, but how deeply models share inner workings.
Alex: That ties the map right back to real limits on transfer. Makes the unevenness feel less random.
Sam: To quantify this, they compared each general model—MSA, multi-dialect, or mixed—to the best dialect-specific one for that region, using a simple gap measure on task scores. Dialect models pulled ahead notably on their home ground, especially for sentiment and word sorting where local word shapes and attachments differ most from formal Arabic. MSA held up well on structure and name spotting across several dialects. The multi-dialect model offered a solid all-around base, often topping MSA on feelings, but it couldn't match top locals even with more training data overall—hinting at clashes between varieties pulling in different directions.
Alex: So the locals win at home turf nuances, but formal carries basics farther.
Sam: Yes. The study notes something called negative interference here—multi-dialect training hurts performance on high-resource dialects compared to focused ones. It challenges the idea that MSA just flows evenly to all dialects; transfer happens, but unevenly by place. MIX models gain an edge on structure and names from including formal text, yet locals win on feelings where daily words matter most.
Alex: Right—like needing a full team's practice hours to outplay veterans, not just a pickup game.
Sam: For inner workings, they checked similarities across model layers in three setups using parallel texts: both on dialect, both on formal, or crossed. Even on matching inputs, alignments stayed below 0.8, meaning models miss each other's fine details. The crossed setup showed weakest links, confirming inputs shape how brains process uniquely. Figures plot dialects by spot on the map from Yemen, bubble size for average task scores or layer matches. Closer ones have bigger bubbles across tasks, fading outward—a clear gradient.
Alex: So function transfers easier than deep structure, even with distance.
Sam: Exactly. On dialect texts, MSA edged structure and names—syntax overlaps, and dialects often borrow formal words for entities. Locals topped feelings, needing grasp of everyday pragmatics. One gap stands out: some dialect models match MSA structure closely yet flop on tasks, showing alignment doesn't guarantee usefulness. The evidence points to distance as a steady predictor for transfer limits within this family.
Alex: Oh—that's why geography predicts behavior but not always the full picture.
Sam: It suggests dialects share bones with formal Arabic but flesh out differently by place. Multi-dialect training hits snags from variety overload on high-resource forms. General models like the MIX one, trained on both formal and dialects, often do better overall than pure dialect ones, thanks to bigger, more varied training sets. But high-resource dialect models, such as those for Egyptian or Saudi Arabic, pull ahead on their home tasks—especially sentiment—while low-resource ones gain from the broader training.
Alex: Fair point. But what about the study's own limits—any catches in how they tested?
Sam: A few. No public dataset exists for Gulf dialects' name spotting, so they used matched sentences and guessed labels with another model—risking some mismatch or errors. Dialect guessing tools aren't perfect, so country splits might have slight mixes. Yemen works as a stand-in for MSA roots from past work, but other spots could shift results. Models vary in size, word lists, and setup, muddling pure comparisons. Some dialect trainings might sneak in formal traces too.
Alex: So noise from data gaps and model differences could tweak the patterns a bit.
Sam: The paper flags these as confounders, suggesting tighter controls—like matched training sizes—would strengthen it. Still, the geography signal holds across methods.
Alex: Makes sense. Practically, for 400 million speakers, this points to smarter designs—like adding tweakable parts for each dialect region to fix gaps without starting over.
Sam: Precisely. Dialect-aware add-ons could smooth transfer, dodging overload from mixing varieties. It questions blanket MSA training, pushing for geography-tuned fixes. That's the grounded takeaway from this work.
Alex: Well said, Sam. This sheds clear light on why AI stumbles on everyday Arabic talk—and how to bridge it. Thanks for breaking it down. Thanks for listening to ResearchPod.