Arabic LLM deployment is often discussed as a language-coverage problem. Coverage matters, but it is not the whole issue. Enterprise Arabic contains dialect, Modern Standard Arabic, transliterated names, English technical terms, legal phrasing, and organization-specific vocabulary.

Evaluation must test meaning, not surface similarity. Did the model preserve the legal force of a clause? Did it distinguish a customer statement from a policy rule? Did it treat a dialectal expression as intent, complaint, instruction, or evidence? Traditional lexical metrics are too blunt for this work.

A defensible path is domain evaluation: test cases from real document patterns, Arabic and bilingual prompts, graded evidence use, and a separation between linguistic fluency and decision correctness.