The best AI for Saudi and Gulf Arabic

Where the models hold Khaleeji dialect and where they slide back into MSA or drift toward Egyptian.

By benchr Editorial Team · · View changelog · Figures verified against official sources, 30 May 2026

The best AI for Saudi and Gulf Arabic: paired scripts and translation paths.
Benchr editorial field plate The best AI for Saudi and Gulf Arabic Writing, translation, and structure
Language modelsThe best AI for Saudi and Gulf Arabic is mapped with paired scripts and translation paths.

One thing to get straight before any picks. This page isn't about translating between Arabic and English. That's a different job with different evaluation criteria, and it has its own guide. This is about register: you write or speak in Khaleeji, and you want the model to answer in Khaleeji rather than defaulting to formal MSA.

That default has a name. Every frontier model is trained on far more Modern Standard Arabic, the formal written register, than on any spoken dialect. So MSA is where they're strong, and dialect is where they wobble. The wobble shows up as drift: you ask in Najdi, and the reply comes back in MSA, or worse, in Egyptian, because Egyptian has more training data than Gulf. Holding the dialect is the whole test here.

Why MSA is the gravity well

Start with the numbers, because they explain everything that follows. On academic Arabic tasks, models sit around 85 to 92 percent. Drop to Gulf dialect and accuracy falls to roughly 75 to 85 percent. Egyptian holds up a little better than Gulf because it's better represented; Levantine sits lower; Maghrebi is the floor. The pattern tracks training-data volume, not anything clever about the dialects themselves.

32% MSA leakage in a Saudi base model before fine-tuning. Specialized LoRA tuning cut it to 6.21%.

That leakage figure is the clearest measurement anyone's published on Gulf drift. A Saudi-tuned model called Saudi-Dialect-ALLaM, trained with Hijazi and Najdi datasets, leaked MSA 32.63 percent of the time before tuning. After LoRA fine-tuning on 5,466 synthetic instruction pairs, leakage dropped to 6.21 percent and its Saudi-dialect rate hit 84.21 percent. The lesson for the frontier models is blunt: none of them are tuned like that, so all of them leak. The question is only how much, and toward what.

Dialect fidelity, model by model

There's no controlled public head-to-head for Claude, GPT, Gemini, and Qwen on Khaleeji Arabic. The shortlist below separates provider documentation from the questions a local review still needs to answer. It is not a measured ranking.

Claude Opus 4.7

Try first Editorial candidate; no public Gulf head-to-head

Qwen 3

Include Broad multilingual positioning; validate Khaleeji

GPT and Gemini

Controls Measure MSA and other-register drift locally

Claude Opus 4.7 is the first candidate in this editorial shortlist, not a verified winner. Anthropic documents multilingual capability generally, but the cited sources do not provide a controlled Gulf-Arabic comparison against GPT, Gemini, or Qwen. An earlier version quoted precise “blind test” scores without a publishable dataset or protocol; those figures have been withdrawn. For a decision, score dialect retention, meaning, tone, and MSA drift separately on your own held-out prompts with fluent reviewers. The related Arabic-content guide is likewise an editorial framework, not a private lab scorecard.

Qwen 3 belongs on the shortlist because Alibaba documents broad multilingual support. Breadth isn't depth: no public benchmark validates Qwen's Khaleeji output against Claude, GPT, or Gemini. If your work is heavily code-mixed, include Qwen and score the switches; for precise Najdi or Hijazi, use fluent local reviewers rather than a language-count claim.

GPT-5.5 should be included when it is already part of your stack, but the cited sources do not establish a Gulf-dialect retention or drift rate. Score whether each response stays in the requested local register, shifts to MSA, or shifts to another dialect; do not infer that behavior from generic Arabic support.

Gemini 3 Pro is another relevant control, especially for teams already using Google products. The cited evidence does not establish how often its output preserves Khaleeji rather than normalizing toward MSA, so measure that explicitly. The broader Gemini 3 Pro evaluation uses the same public-evidence boundary.

Najdi and Hijazi: the sub-dialect cliff

Step inside Saudi Arabia and the ground gets thinner. "Gulf Arabic" in a benchmark usually means a blended Khaleeji average. Najdi (central, Riyadh) and Hijazi (western, Jeddah and Mecca) are distinct, and they're rare enough in public data that no frontier model can promise either one. The literature is consistent here: large models stay dominated by MSA with limited support for Saudi dialects specifically, and they tend to collapse Najdi and Hijazi into a generic Gulf register or straight into MSA.

A model can know "Gulf" as a category and still flatten Najdi and Hijazi into the same beige Arabic.

This is where fine-tuned models may earn their place. If your product depends on Saudi sub-dialect accuracy, the published Saudi-Dialect-ALLaM research is a useful proof of concept for specialization. The cited public evidence does not establish which off-the-shelf frontier model collapses sub-dialects least, so test that question directly rather than treating this editorial shortlist as a result.

Evaluation shortlist

Use these as candidates to test, not measured winners. Keep the same held-out prompts and blind rubric across the shortlist.

General Khaleeji chat

Start with Claude Editorial candidate; measure register drift

Arabic-English code-mix

Include Qwen Broad support claim; score every switch

Najdi / Hijazi accuracy

Add a specialist Where licensing and deployment allow

Formal MSA output

Test all Generic Arabic support does not rank style

Voice / transcription

Speechmatics Publicly reported code-switch WER; verify corpus fit

Drift controls

GPT + Gemini Measure MSA and other-dialect shifts

A note on that voice row, because it's easy to over-read. Speechmatics' Arabic-English bilingual model hits a 6.3 percent word error rate on code-switching, about 35 percent lower than Google's 9.7 percent. That's a strong result, but it's speech-to-text. It tells you nothing about how a chat model generates dialect, so don't carry it over to text work. If you're weighing the spoken side more broadly, benchr's comparison of voice models has the wider picture.

If your real need is moving text cleanly between the two languages rather than holding one dialect, that's the sibling guide's territory: benchr's guide to Arabic-English translation ranks the models on direction quality both ways, which is a separate question from register fidelity. And if the underlying job is just producing solid long-form Arabic prose, the language-agnostic guide to AI for writing is the better starting point, since the writing-quality leader and the dialect leader happen to be the same family.

Calculate your cost →·Compare this model →·Find your model →

Frequently asked

Does Claude handle Khaleeji (Gulf Arabic) better than other models?

No controlled public head-to-head establishes that result. Claude is this page's editorial first candidate based on broader multilingual documentation and public reporting, not a benchr test or a published Gulf score. Compare it directly on your held-out prompts.

What is the difference between MSA and Gulf dialect performance in LLMs?

Models perform much better on Modern Standard Arabic, around 85 to 92 percent on academic benchmarks, than on Gulf dialect, around 75 to 85 percent, because training data skews toward formal written Arabic. When a model hits Gulf input it often drifts back to MSA or a more common dialect, and the effect is worst on underrepresented sub-dialects like Najdi and Hijazi Saudi Arabic.

How do sub-dialects like Najdi and Hijazi perform in modern LLMs?

Najdi and Hijazi are badly underrepresented in every major frontier model. Base models like Claude, GPT, Qwen, and Gemini do not explicitly separate or optimize for these sub-dialects and tend to collapse them toward MSA or a generic Gulf register. Specialized fine-tuning closes the gap: Saudi-Dialect-ALLaM, LoRA-tuned on 5,466 synthetic instruction pairs, reached an 84.21 percent Saudi rate and cut MSA leakage from 32.63 percent to 6.21 percent.

Does Qwen 3 compete with Claude on Arabic dialects?

Alibaba documents broad multilingual support for Qwen, which makes it a reasonable candidate to include. No published controlled benchmark establishes that it handles Gulf Arabic or Arabic-English code-switching better than Claude, GPT, or Gemini. Test all candidates on the same held-out local prompts; multilingual breadth does not guarantee Gulf depth.

Which model best handles Arabic-English code-switching?

The cited Speechmatics result compares speech transcription on one Arabic-English code-switching corpus; it does not establish performance for text-generating chat models. No controlled public comparison identifies a best frontier text model here. Build a held-out set of the exact switches your users make and score meaning, register, and unwanted normalization with fluent reviewers.

Use fixed Gulf-Arabic cases and qualified reviewers

The Gulf pack publishes ten fixed prompts for MSA, Saudi support, dialect ambiguity, code-switching, privacy, translation, and RTL formatting. It does not publish a universal model winner.

  • Blind model names before native reviewers score dialect and service tone.
  • Keep technical English where Gulf teams normally keep it.
  • Record reviewer disagreement instead of hiding it inside one average.

Changelog

  • August 21, 2026 — Added evaluation pack arabic-gulf-v1 with inspectable prompts and rubrics; no unmeasured model result is published.
  • July 23, 2026 — Withdrew unsupported blind-test scores and recast the model order as an explicitly editorial shortlist requiring local Gulf-Arabic validation.
  • May 30, 2026 — Originally published. Covers Claude Opus 4.7, Qwen 3, GPT-5.5, and Gemini 3 Pro on Gulf register and MSA drift; sub-dialect and fine-tuning notes from current research.

References

  1. Truescho, "Claude vs ChatGPT: Which Is Better for Arabic Content? (2026)," truescho.com, accessed May 2026.
  2. Anthropic, "Multilingual support," platform.claude.com, accessed May 2026.
  3. "Saudi-Dialect-ALLaM: LoRA Fine-Tuning for Dialectal Arabic Generation," arxiv.org, accessed May 2026.
  4. "Advancing AI-Driven Linguistic Analysis: Arabic Dialect Corpora for Gulf Countries and Saudi Arabia," mdpi.com, accessed May 2026.
  5. "Cross-dialectal Arabic translation: comparative analysis on large language models," frontiersin.org, accessed May 2026.
  6. Speechmatics, "Arabic-English bilingual speech-to-text," speechmatics.com, accessed May 2026.