One thing to get straight before any picks. This page isn't about translating between Arabic and English. That's a different job with different evaluation criteria, and it has its own guide. This is about register: you write or speak in Khaleeji, and you want the model to answer in Khaleeji rather than defaulting to formal MSA.
That default has a name. Every frontier model is trained on far more Modern Standard Arabic, the formal written register, than on any spoken dialect. So MSA is where they're strong, and dialect is where they wobble. The wobble shows up as drift: you ask in Najdi, and the reply comes back in MSA, or worse, in Egyptian, because Egyptian has more training data than Gulf. Holding the dialect is the whole test here.
Why MSA is the gravity well
Start with the numbers, because they explain everything that follows. On academic Arabic tasks, models sit around 85 to 92 percent. Drop to Gulf dialect and accuracy falls to roughly 75 to 85 percent. Egyptian holds up a little better than Gulf because it's better represented; Levantine sits lower; Maghrebi is the floor. The pattern tracks training-data volume, not anything clever about the dialects themselves.
That leakage figure is the clearest measurement anyone's published on Gulf drift. A Saudi-tuned model called Saudi-Dialect-ALLaM, trained with Hijazi and Najdi datasets, leaked MSA 32.63 percent of the time before tuning. After LoRA fine-tuning on 5,466 synthetic instruction pairs, leakage dropped to 6.21 percent and its Saudi-dialect rate hit 84.21 percent. The lesson for the frontier models is blunt: none of them are tuned like that, so all of them leak. The question is only how much, and toward what.
Dialect fidelity, model by model
There's no controlled public head-to-head for Claude, GPT, Gemini, and Qwen on Khaleeji Arabic. The shortlist below separates provider documentation from the questions a local review still needs to answer. It is not a measured ranking.
Claude Opus 4.7
Try first Editorial candidate; no public Gulf head-to-headQwen 3
Include Broad multilingual positioning; validate KhaleejiGPT and Gemini
Controls Measure MSA and other-register drift locallyClaude Opus 4.7 is the first candidate in this editorial shortlist, not a verified winner. Anthropic documents multilingual capability generally, but the cited sources do not provide a controlled Gulf-Arabic comparison against GPT, Gemini, or Qwen. An earlier version quoted precise “blind test” scores without a publishable dataset or protocol; those figures have been withdrawn. For a decision, score dialect retention, meaning, tone, and MSA drift separately on your own held-out prompts with fluent reviewers. The related Arabic-content guide is likewise an editorial framework, not a private lab scorecard.
Qwen 3 belongs on the shortlist because Alibaba documents broad multilingual support. Breadth isn't depth: no public benchmark validates Qwen's Khaleeji output against Claude, GPT, or Gemini. If your work is heavily code-mixed, include Qwen and score the switches; for precise Najdi or Hijazi, use fluent local reviewers rather than a language-count claim.
GPT-5.5 should be included when it is already part of your stack, but the cited sources do not establish a Gulf-dialect retention or drift rate. Score whether each response stays in the requested local register, shifts to MSA, or shifts to another dialect; do not infer that behavior from generic Arabic support.
Gemini 3 Pro is another relevant control, especially for teams already using Google products. The cited evidence does not establish how often its output preserves Khaleeji rather than normalizing toward MSA, so measure that explicitly. The broader Gemini 3 Pro evaluation uses the same public-evidence boundary.
Najdi and Hijazi: the sub-dialect cliff
Step inside Saudi Arabia and the ground gets thinner. "Gulf Arabic" in a benchmark usually means a blended Khaleeji average. Najdi (central, Riyadh) and Hijazi (western, Jeddah and Mecca) are distinct, and they're rare enough in public data that no frontier model can promise either one. The literature is consistent here: large models stay dominated by MSA with limited support for Saudi dialects specifically, and they tend to collapse Najdi and Hijazi into a generic Gulf register or straight into MSA.
A model can know "Gulf" as a category and still flatten Najdi and Hijazi into the same beige Arabic.
This is where fine-tuned models may earn their place. If your product depends on Saudi sub-dialect accuracy, the published Saudi-Dialect-ALLaM research is a useful proof of concept for specialization. The cited public evidence does not establish which off-the-shelf frontier model collapses sub-dialects least, so test that question directly rather than treating this editorial shortlist as a result.
Evaluation shortlist
Use these as candidates to test, not measured winners. Keep the same held-out prompts and blind rubric across the shortlist.
General Khaleeji chat
Start with Claude Editorial candidate; measure register driftArabic-English code-mix
Include Qwen Broad support claim; score every switchNajdi / Hijazi accuracy
Add a specialist Where licensing and deployment allowFormal MSA output
Test all Generic Arabic support does not rank styleVoice / transcription
Speechmatics Publicly reported code-switch WER; verify corpus fitDrift controls
GPT + Gemini Measure MSA and other-dialect shiftsA note on that voice row, because it's easy to over-read. Speechmatics' Arabic-English bilingual model hits a 6.3 percent word error rate on code-switching, about 35 percent lower than Google's 9.7 percent. That's a strong result, but it's speech-to-text. It tells you nothing about how a chat model generates dialect, so don't carry it over to text work. If you're weighing the spoken side more broadly, benchr's comparison of voice models has the wider picture.
If your real need is moving text cleanly between the two languages rather than holding one dialect, that's the sibling guide's territory: benchr's guide to Arabic-English translation ranks the models on direction quality both ways, which is a separate question from register fidelity. And if the underlying job is just producing solid long-form Arabic prose, the language-agnostic guide to AI for writing is the better starting point, since the writing-quality leader and the dialect leader happen to be the same family.
Calculate your cost →·Compare this model →·Find your model →