Translation direction, register, audience, and domain are separate variables. A short MSA email, a Gulf customer-service transcript, a contract, and a classical text should not share one score. Public model specifications can identify candidates, but they do not establish dialect quality or a winner for your material.
Editorial candidates to test
The shortlist below is a set of editorial hypotheses, not a published comparative result. Hosted frontier models are candidates when their documented context and workflow features fit the job. Qwen 3 and Llama 4 are additional candidates when open weights, local hosting, licensing, or runtime control matter. A result reported for one translation-tuned variant cannot be transferred to an entire model family.
| Candidate | Documented reason to shortlist | What the local test must decide |
|---|---|---|
| Claude Opus 4.8 | Hosted model with a documented long-context workflow | Meaning, register, terminology, names, consistency, latency, and cost on your files |
| GPT-5.5 | Hosted model and API workflow documented by its provider | Both directions and every required regional register |
| Gemini 3.1 Pro | Hosted model with documented long-context and Google workflow options | Instruction adherence, source coverage, register, and current plan fit |
| Qwen 3 / exact translation variant | Open-weight or translation-specific deployment may offer more runtime control | The exact checkpoint, tokenizer, serving stack, and license you will deploy |
| Llama 4 / exact checkpoint | Open-weight deployment may fit local-hosting or governance requirements | The exact checkpoint and runtime; no family-wide Arabic rank is assumed |
A reproducible blind evaluation
- Freeze the setup. Record the exact model ID, date, system prompt, temperature, tools, glossary, and output constraints. Use the same instructions and source text for every candidate.
- Stratify the set. Include both translation directions and the registers you will ship: MSA, Gulf, Egyptian, Levantine, Maghrebi, code-switched text, legal terminology, names and transliteration, and classical or religious material when relevant.
- Blind the review. Randomize and relabel outputs so native reviewers do not know the provider. Use more than one reviewer for subjective register judgments.
- Score defined criteria. Measure preservation of meaning, omissions and additions, terminology, register, names, formatting, and reviewer preference. Track disagreement instead of collapsing it into an unsupported winner claim.
- Repeat and report. Rerun a sample to check stability. Publish the test set, prompt, model IDs, scoring rubric, mean scores, reviewer disagreement, latency, and cost separately.
Required stress tests
State the intended audience, country, register, glossary, and transliteration convention explicitly. For legal, medical, religious, or other consequential text, model preference is not a substitute for domain review. If a required dialect has little written standardization, document reviewer disagreement and test real examples from that audience instead of extrapolating from MSA benchmarks.
Context-window size is also not evidence of long-document translation quality. Test full-document terminology, cross-reference preservation, omissions, and recovery from truncation on representative files. If the document exceeds a system's practical limits, use a controlled segmentation and glossary process and verify consistency after recombination.
Calculate your cost →·Compare this model →·Find your model →