Fluency alone does not prove that the Arabic works. A Khaleeji reviewer may catch Egyptian markers, the wrong politeness level, or a product term that should have stayed in English. That is why this article works as an evaluation sheet: official product pages define the shortlist, while the tasks and review criteria give you a repeatable test. Forum discussion is useful context, not a controlled benchmark.
The candidate set is Claude, GPT, Gemini, Qwen, and Llama. The workload axes are Modern Standard Arabic, three regional registers (Khaleeji, Egyptian, Levantine), and code-switched Arabic-English customer email. Maghrebi Arabic (Moroccan, Algerian, Tunisian) is a separate evaluation and is not in scope here.
The Arabic evaluation worksheet
Run every row with the same model snapshot, system instructions, sampling settings, and source text. Randomize model labels before review. Use at least two fluent reviewers for customer-facing dialect work, record disagreements, and keep the prompts and raw outputs with the result.
| Workload | Prompt control | Pass criteria | Required review |
|---|---|---|---|
| EN → Khaleeji marketing | Name city, age, channel, and tone | No register drift; product terms preserved | Two local copy reviewers |
| MSA business reply | Fix audience and formality | Accurate facts and suitable politeness | Business-language editor |
| Literary paraphrase | Provide a lawful reference excerpt | Meaning retained; no invented quotation | Arabic editor |
| Labor-law summary (MSA) | Supply authoritative source text | Every legal claim traceable to source | Qualified legal reviewer |
| Egyptian → English | Include idioms and implied meaning | No omission or literalized idiom | Bilingual reviewer |
| Code-switched email | Mark terms that must stay in English | Natural register and term retention | Target-market reviewer |
Khaleeji tests must be local
A useful Khaleeji test is English-to-Arabic marketing copy for a precisely defined Saudi audience. Give every candidate the same brand terms, city, age range, channel, and tone. Ask reviewers to mark dialect substitutions, unwanted MSA, non-Saudi phrasing, unnatural rhythm, and any brand term that should have remained in Latin script.
Do not encode an expected winner into the rubric. Shuffle the outputs and let reviewers label observed drift without seeing the model name. If the reviewers disagree, preserve that disagreement instead of converting it into a precise score. Re-run the test when a model snapshot or system prompt changes.
MSA still needs task-specific checks
MSA fluency does not establish factual accuracy or fitness for a particular audience. Score terminology, sentence structure, tone, omissions, and unsupported additions separately. A polished paragraph can still fail the task.
For a Saudi-business support reply, give each candidate the same customer history and approved terminology. Reviewers should score whether the opening and close fit the relationship, whether the answer resolves the issue, and whether any sentence sounds translated or excessively formal. Those observations become evidence for your deployment; they are not transferable as a universal model ranking.
Fluency is only the first gate. Tone, rhythm, local idiom, factual accuracy, and review effort decide whether the text is usable.
Register
Blind scoreMSA, Khaleeji, Egyptian, LevantineTerminology
Lock termsNames and technical vocabularyMeaning
Check sourceOmissions and unsupported additionsTone
Local reviewAudience and relationship fitSafety
Expert gateLegal, medical, religious textOperations
Track editsTime and rework per candidateCode-switching is a useful stress test
Mixed Arabic-English customer email tests several controls at once. Saudi internet writing may keep brand names and technical terms in English inside otherwise-Arabic prose. Define which terms must stay unchanged and what relationship and register the reply should convey.
Score over-translation, under-translation, term corruption, and register mismatch as separate failure types. Use several real, de-identified examples and a native reviewer; a single polished output is not enough to establish a product-wide conclusion.
MSA
Terminology Accuracy and formalityKhaleeji
Locale City and audience fitEgyptian
Idioms Meaning, not literal wordsLevantine
Register No cross-dialect driftCode-switching
Term retention Arabic-English boundariesPoetry
Meaning No invented quotationMSA, Khaleeji, Egyptian, or Levantine. Define the target before testing.
Filter by privacy, deployment, budget, and the exact license you can use.
Name the city, the age range, the tone. Defaults aren't enough.
For customer-facing copy, always. Don't skip this step.
Where model output needs an expert gate
Some categories carry enough consequence that fluency is not an acceptable release criterion. Use model output only as a draft and require an accountable expert before publication.
Legal text. MSA legal summaries are good enough to draft from but not to publish. Specific terms carry specific meanings; misremembering an article number or substituting a near-synonym changes the legal implication. Don't deploy any of these models for Arabic legal work without a qualified human reviewer.
Classical Arabic. Check quotations from medieval texts, religious exegesis, and classical-style passages against primary sources. Require a qualified specialist to review any translation or interpretation.
Specific regional dialects. Khaleeji is itself a family of varieties: Najdi, Hijazi, Qatari, and Bahraini are not interchangeable. Name the target precisely and recruit reviewers from that audience instead of treating “Khaleeji” as one score.
How to make the production choice
Start with candidates that satisfy your privacy, hosting, latency, and budget constraints. Include both hosted and open-weight options when your deployment permits. Product documentation can establish availability and configuration; only your held-out Arabic set can establish fit for your audience.
If an open-weight model such as Qwen is on the shortlist, verify the exact model card and license for the checkpoint you will deploy. The open-weight tier piece explains deployment tradeoffs, and the small-models piece covers when a smaller model may be sufficient.
Do not carry a verdict from general MSA into dialect writing, document images, or translation. Treat each as a separate workload with its own source set and pass criteria. The multimodal piece covers a separate way to evaluate document-image tasks.
For customer-facing Saudi work, retain a Khaleeji-fluent reviewer regardless of the selected model. Track acceptance rate, factual corrections, dialect corrections, and editing time by task. Re-evaluate after material model or prompt changes; that measured workflow is more defensible than a permanent brand ranking.