Many RAG programs begin with a model comparison and only later ask whether retrieval is producing the right evidence. That sequence is backwards. If the retrieval layer returns the wrong clause, duplicated passages, stale policies, or context fragments without hierarchy, the generator is being asked to reason over a corrupted record.

The retrieval episode, not only the final answer, should become the unit of evaluation. Which document was selected? Which section? Was the answerable span present? Did the system retrieve redundant text at the expense of decisive evidence? Did retrieval preserve jurisdiction, date, policy version, and exception language?

In Arabic and bilingual enterprise settings, similarity is a weak proxy when morphology, dialect, boilerplate, and translated policy language interact. The remedy is an evidence discipline: retrieval test sets, gold spans, citation faithfulness, and context assembly tuned before a larger model is celebrated.