When Size Doesn't Buy You Dialect
Arabic isn't one language—it's a family pretending to be one. Modern Standard Arabic dominates formal text, but day-to-day conversation in the UAE happens in Emirati Arabic, a Gulf dialect with its own vocabulary, rhythm, and cultural freight that MSA just doesn't carry.
Falcon-Emirati-7B is TII's new dialect-specialized chat model, built on top of their existing Falcon-H1-Arabic foundation. It scores 84.83% on Alyah, a native Emirati-dialect benchmark, beating every other Arabic and multilingual model tested—including several many times its size. That performance gap tells you something important: dialect competence doesn't emerge from scale. You have to train for it on purpose.
The really interesting part isn't the benchmark number. It's how they got there, and what broke along the way.
Starting From Falcon-H1-Arabic
TII didn't start from scratch. Falcon-Emirati-7B builds on Falcon-H1-Arabic, their hybrid-architecture Arabic model family released earlier this year. The architecture runs State Space Models (Mamba) and Transformer attention in parallel inside every block, fusing outputs before projection—linear-time efficiency on long sequences plus the precision of attention for long-range dependencies.
The H1-Arabic family spans 3B, 7B, and 34B parameters with context windows up to 128K and 256K tokens, pretrained on a broad mix of MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) alongside English and multilingual data. That gave them a strong starting point: a model that already understood Arabic broadly and had some dialectal exposure baked in.
They picked the 7B variant specifically. Large enough to hold onto the nuance dialect adaptation needs, small enough that training and inference stay practical. The 34B would push quality further but at a cost that doesn't make sense for a dialect-specialized chat model. The 3B doesn't leave enough headroom for the depth of cultural and linguistic understanding they were after.
Why Dialect Adaptation Is Genuinely Hard
Turning a general Arabic model into an Emirati specialist sounds easier than building the base model. It isn't, for three reasons:
- Emirati is mostly spoken. It shows up far less in writing online than MSA or even other Gulf and Levantine dialects, so there just isn't as much raw text to learn from.
- Meaning is often non-literal. Idioms, proverbs, and poetic references (especially nabati poetry) lean on shared cultural context, not surface vocabulary. A model that only knows MSA can translate every word of an Emirati sentence and still miss what it actually means.
- There's no established playbook. There isn't a documented recipe for how much dialectal data is enough, how to mix it with MSA, or which training stage—continued pretraining, SFT, or preference optimization—matters most for picking up a dialect.
That last point shaped how they worked. A lot of building Falcon-Emirati came down to trial and error: testing different data mixes, training stages, and supervision strategies, then using both human judgment and benchmark scores to figure out what actually moved the needle.
The Data Pipeline: Three Complementary Sources
TII built a dedicated Emirati data pipeline drawing on three sources:
1. Authentic Emirati-Dialect Web Data
They crawled and curated content from Emirati websites and forums written natively in the dialect—not translated or transliterated from MSA. This is ground truth: how Emiratis actually write and speak online, the everyday phrasing, the colloquial expressions, and the natural back-and-forth between Emirati and MSA that shows up in real usage.
2. MSA Data About Emirati Culture and Identity
Alongside dialectal text, they pulled in MSA-language material specifically about Emirati culture, heritage, and language: articles on local customs, values, history, social norms, including how Emiratis are perceived and stereotyped.
This doesn't teach the model to write in dialect, but it teaches the model what it's talking about when Emirati topics come up—heritage, etiquette, context a native speaker just knows.
3. Synthetic Data, Guided by Glossaries and Style Rules
Authentic dialectal text alone wasn't enough to cover the range of topics a chat model needs to handle day to day. So they generated synthetic Emirati-dialect data to fill gaps.
They didn't just let a generator improvise in "Gulf-ish" Arabic. They constrained it with strict rules, glossaries, and dictionaries built specifically for Emirati vocabulary and grammar. Those guardrails made the difference between synthetic output that reads as authentically Emirati and output that's grammatically fine but sounds off to anyone who actually speaks the dialect.
Finding the Recipe Through Ablations
Since there's no standard recipe for MSA-to-dialect adaptation, they treated the training strategy itself as an experimental question. They ran ablations on:
- How much dialectal data to inject and at which stage of training
- How to balance authentic crawled data against synthetic data without overfitting to synthetic patterns
- How much MSA cultural context was actually needed to keep the model culturally grounded rather than just fluent on the surface
At each step they leaned on a mix of automatic scoring and native-speaker review, since automatic metrics alone don't capture naturalness, tone, or cultural fit well enough to trust.
Evaluation: Benchmarks Plus Native Ears
They tracked progress with two complementary approaches. First, native Emirati speakers reviewed outputs directly, judging not just correctness but whether answers sounded right: naturalness, tone, cultural appropriateness. The things a benchmark score won't tell you but a native ear catches immediately.
Second, they used Alyah (الياه, "North Star"), a benchmark they and the community released specifically to evaluate Emirati-dialect capability. Alyah is a fully native multiple-choice benchmark of 1,173 samples, collected manually from native Emirati speakers and spanning categories from everyday greetings and etiquette to figurative language, heritage knowledge, and Emirati poetry—the categories where dialect and culture matter most and where generic Arabic models tend to struggle.
The Numbers: 84.83% on Alyah
Falcon-Emirati-7B scores 84.83% on Alyah, ahead of every other Arabic and multilingual model they compared it against. A couple of things jump out:
Size alone doesn't buy you dialect competence. Some of the largest multilingual models score well below smaller, more dialect-aware ones. Emirati proficiency has to be trained for on purpose, not picked up as a side effect of scale.
The models that do best tend to be Arabic-native or Arabic-focused to begin with. General Arabic and dialect coverage is a necessary starting point, but it still takes targeted, dialect-specific work to close the rest of the gap, particularly on the hardest parts of Alyah: poetry, heritage knowledge, and the language-and-dialect category itself.
This matches what came out of the Alyah benchmark release more broadly. Even strong models show real degradation once you move into genuinely dialectal, culturally embedded content. That gap doesn't close on its own with bigger models. It takes data and evaluation built specifically for the dialect.
Beyond Multiple Choice: Does It Actually Generate Emirati?
Multiple-choice accuracy tells you whether a model can recognize the right answer among four options. It doesn't tell you whether the model will produce Emirati Arabic on its own when someone just talks to it.
So they ran a second evaluation: open-ended generation on the same 1,173 Alyah questions, scored by an LLM judge (Gemini 3.7 Flash) against five models. The judge scored each answer on two separate dimensions: whether the content was correct, and independently, whether the answer actually came back in Emirati dialect rather than MSA.
Falcon-Emirati-7B leads on correctness, but the real gap is in dialect fidelity. It scores 0.52 (partial credit) against 0.05 for ALLaM-7B-Instruct-preview, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat, and effectively 0.00 for Fanar-2-27B-Instruct.
That's not a small edge. It's close to two orders of magnitude at the low end. In practice, the other models often know the right answer but default to MSA when generating. They can recognize Emirati, but they don't speak it back.
What This Tells Us About Dialect Models
The Falcon-Emirati work surfaces a few things worth generalizing:
Dialect adaptation is a distinct capability from general language modeling. You can't assume a strong Arabic foundation will automatically extend to fluent, culturally appropriate Emirati generation. The same is likely true for other low-resource dialects and language varieties.
Synthetic data works, but only with strong constraints. Unconstrained generation in "dialect-ish" language sounds off to native speakers. Glossaries, style rules, and strict vocabularies turn synthetic data from filler into a useful training signal.
Evaluation has to measure generation, not just recognition. Multiple-choice benchmarks are useful for tracking, but they don't tell you whether your model will actually produce the dialect when users expect it to. You need open-ended generation scored by judges who understand the dialect, whether human or LLM.
Cultural grounding matters as much as linguistic fidelity. Training on MSA material about Emirati culture, alongside dialectal text, helped the model understand what it was talking about, not just how to say it. That distinction shows up in human eval even when benchmarks don't catch it.
TII built a model that actually speaks Emirati, not just a model that scores well on Emirati questions. That required intention, ablations, and a willingness to treat the recipe itself as an open research question. The gap between the two—recognition versus generation—is where a lot of dialect work still lives.