AI for Arabic content: a working report on five models

How Claude, GPT-5, Gemini, Qwen, and Llama handle MSA and three regional Arabic dialects, and where every model still struggles.

By benchr Editorial Team · · View changelog

AI for Arabic content: a working report on five models: paired scripts and translation paths.
Benchr editorial field plate AI for Arabic content Writing, translation, and structure
Language modelsThe visual for AI for Arabic content: a working report on five models pairs paired scripts and translation paths.
Models compared 5 Claude, GPT-5, Gemini, Qwen, Llama
Workload axes 6 Across MSA and three dialects
Saudi-market gate Native review Blind score before publishing
Open-weight gate License check Verify the exact model card

Fluency alone does not prove that the Arabic works. A Khaleeji reviewer may catch Egyptian markers, the wrong politeness level, or a product term that should have stayed in English. That is why this article works as an evaluation sheet: official product pages define the shortlist, while the tasks and review criteria give you a repeatable test. Forum discussion is useful context, not a controlled benchmark.

The candidate set is Claude, GPT, Gemini, Qwen, and Llama. The workload axes are Modern Standard Arabic, three regional registers (Khaleeji, Egyptian, Levantine), and code-switched Arabic-English customer email. Maghrebi Arabic (Moroccan, Algerian, Tunisian) is a separate evaluation and is not in scope here.

The Arabic evaluation worksheet

Run every row with the same model snapshot, system instructions, sampling settings, and source text. Randomize model labels before review. Use at least two fluent reviewers for customer-facing dialect work, record disagreements, and keep the prompts and raw outputs with the result.

Reusable test matrix for Saudi-market Arabic workloads
WorkloadPrompt controlPass criteriaRequired review
EN → Khaleeji marketingName city, age, channel, and toneNo register drift; product terms preservedTwo local copy reviewers
MSA business replyFix audience and formalityAccurate facts and suitable politenessBusiness-language editor
Literary paraphraseProvide a lawful reference excerptMeaning retained; no invented quotationArabic editor
Labor-law summary (MSA)Supply authoritative source textEvery legal claim traceable to sourceQualified legal reviewer
Egyptian → EnglishInclude idioms and implied meaningNo omission or literalized idiomBilingual reviewer
Code-switched emailMark terms that must stay in EnglishNatural register and term retentionTarget-market reviewer

Khaleeji tests must be local

A useful Khaleeji test is English-to-Arabic marketing copy for a precisely defined Saudi audience. Give every candidate the same brand terms, city, age range, channel, and tone. Ask reviewers to mark dialect substitutions, unwanted MSA, non-Saudi phrasing, unnatural rhythm, and any brand term that should have remained in Latin script.

Do not encode an expected winner into the rubric. Shuffle the outputs and let reviewers label observed drift without seeing the model name. If the reviewers disagree, preserve that disagreement instead of converting it into a precise score. Re-run the test when a model snapshot or system prompt changes.

MSA still needs task-specific checks

MSA fluency does not establish factual accuracy or fitness for a particular audience. Score terminology, sentence structure, tone, omissions, and unsupported additions separately. A polished paragraph can still fail the task.

For a Saudi-business support reply, give each candidate the same customer history and approved terminology. Reviewers should score whether the opening and close fit the relationship, whether the answer resolves the issue, and whether any sentence sounds translated or excessively formal. Those observations become evidence for your deployment; they are not transferable as a universal model ranking.

Fluency is only the first gate. Tone, rhythm, local idiom, factual accuracy, and review effort decide whether the text is usable.

Register

Blind scoreMSA, Khaleeji, Egyptian, Levantine

Terminology

Lock termsNames and technical vocabulary

Meaning

Check sourceOmissions and unsupported additions

Tone

Local reviewAudience and relationship fit

Safety

Expert gateLegal, medical, religious text

Operations

Track editsTime and rework per candidate

Code-switching is a useful stress test

Mixed Arabic-English customer email tests several controls at once. Saudi internet writing may keep brand names and technical terms in English inside otherwise-Arabic prose. Define which terms must stay unchanged and what relationship and register the reply should convey.

Score over-translation, under-translation, term corruption, and register mismatch as separate failure types. Use several real, de-identified examples and a native reviewer; a single polished output is not enough to establish a product-wide conclusion.

MSA

Terminology Accuracy and formality

Khaleeji

Locale City and audience fit

Egyptian

Idioms Meaning, not literal words

Levantine

Register No cross-dialect drift

Code-switching

Term retention Arabic-English boundaries

Poetry

Meaning No invented quotation
1. Identify the register

MSA, Khaleeji, Egyptian, or Levantine. Define the target before testing.

2. Build the shortlist

Filter by privacy, deployment, budget, and the exact license you can use.

3. Prompt with audience

Name the city, the age range, the tone. Defaults aren't enough.

4. Native-speaker review

For customer-facing copy, always. Don't skip this step.

Where model output needs an expert gate

Some categories carry enough consequence that fluency is not an acceptable release criterion. Use model output only as a draft and require an accountable expert before publication.

Legal text. MSA legal summaries are good enough to draft from but not to publish. Specific terms carry specific meanings; misremembering an article number or substituting a near-synonym changes the legal implication. Don't deploy any of these models for Arabic legal work without a qualified human reviewer.

Classical Arabic. Check quotations from medieval texts, religious exegesis, and classical-style passages against primary sources. Require a qualified specialist to review any translation or interpretation.

Specific regional dialects. Khaleeji is itself a family of varieties: Najdi, Hijazi, Qatari, and Bahraini are not interchangeable. Name the target precisely and recruit reviewers from that audience instead of treating “Khaleeji” as one score.

How to make the production choice

Start with candidates that satisfy your privacy, hosting, latency, and budget constraints. Include both hosted and open-weight options when your deployment permits. Product documentation can establish availability and configuration; only your held-out Arabic set can establish fit for your audience.

If an open-weight model such as Qwen is on the shortlist, verify the exact model card and license for the checkpoint you will deploy. The open-weight tier piece explains deployment tradeoffs, and the small-models piece covers when a smaller model may be sufficient.

Do not carry a verdict from general MSA into dialect writing, document images, or translation. Treat each as a separate workload with its own source set and pass criteria. The multimodal piece covers a separate way to evaluate document-image tasks.

For customer-facing Saudi work, retain a Khaleeji-fluent reviewer regardless of the selected model. Track acceptance rate, factual corrections, dialect corrections, and editing time by task. Re-evaluate after material model or prompt changes; that measured workflow is more defensible than a permanent brand ranking.

Frequently asked

Which AI model is best for Arabic content?

There is no public controlled benchmark that establishes one winner for Saudi-market Arabic. Shortlist models that meet your deployment and licensing needs, then blind-score them on held-out MSA, Khaleeji, and code-switched prompts with native reviewers.

Can AI write in Saudi (Khaleeji) Arabic?

Yes, as a draft. Specify the city or regional variety, audience, channel, and tone; then have a fluent local reviewer check register, idiom, and unintended drift before publication.

Does Qwen 3 handle Arabic well?

Qwen is a reasonable candidate to include in an Arabic evaluation, especially when an open-weight deployment matters. Verify the exact model card and license, then test its dialect fit on your own prompts rather than assuming a public ranking.

How well does AI handle code-switched Arabic-English?

Treat mixed Arabic-English email as a useful stress test, not a proven hardest task. Score term preservation, requested register, and incorrect translation of names or technical vocabulary.

Can AI translate medieval or Classical Arabic?

Treat pre-modern and religious text as high-risk. Check every quotation against a primary source and require review by a qualified specialist before publishing a translation or interpretation.

Changelog

  • July 23, 2026 — Removed unsupported model rankings and community-consensus claims. Reframed the comparison as a blind, repeatable Arabic evaluation worksheet with explicit human-review gates.
  • May 25, 2026 — Pre-publication draft: replaced an earlier private scoring narrative with an initial qualitative comparison.
  • May 30, 2026 — Published with retrospective coverage through the May 4, 2026 subject date.

References

  1. Anthropic, "Claude API Documentation," docs.claude.com, accessed May 2026.
  2. Alibaba, "Qwen," qwen.ai, accessed May 2026.
  3. Google, "Gemini API models," ai.google.dev/gemini-api/docs/models, accessed May 2026.
  4. Meta, "Llama," llama.com, accessed May 2026.
  5. "Chatbot Arena leaderboard," lmarena.ai, May 2026 snapshot.
  6. "Modern Standard Arabic," Wikipedia, en.wikipedia.org/wiki/Modern_Standard_Arabic, accessed May 2026.