A fluent Arabic answer can still be operationally wrong. That is the central evaluation problem for Arabic LLMs in enterprise settings. The issue is not only whether the model speaks Arabic, but whether it understands the context in which Arabic is being used.
Context includes role, policy source, jurisdiction, document type, dialect, and implied intent. A claims note, a regulatory circular, and a customer message may share vocabulary while requiring different reasoning. Evaluation must make these distinctions visible.
For governance, Arabic evaluation should include grounded answer tests, refusal tests, translation-consistency checks, and domain-specific adjudication. The goal is not to prove that a model is generally good at Arabic. The goal is to know where it is safe to use.