Essays
Production AI, retrieval, agents, and Arabic-English enterprise systems.
A technical archive on evaluation, governance, tool use, retrieval discipline, and the institutional controls around AI.
RPA for Brokers: Control Before Automation
Broker automation needs a control layer, not another set of screens wired together.
Fine-Tuning FLOPs as an AI Governance Signal
Compute accounting is becoming a governance record, not only an infrastructure metric.
MCP and the Future of Context Management
MCP should be understood as an interface discipline: it makes context and tools explicit enough to govern.
What Teams Get Wrong About MCPs in Production
MCP is useful only when teams treat context, tools, and session state as governed interfaces.
MCP Agents Need Session Integrity
The hard problem is not calling one tool; it is preserving meaning across a sequence of stateful tool interactions.
Why MCP Matters for Agent Infrastructure
MCP turns ad hoc agent integrations into contracts that can be tested, governed, and changed.
Gulf Enterprise Agents Need Closed-Loop Control
For Gulf enterprise agents, reliability depends on validating actions before they touch sensitive systems.
Fann or Flop: AI Agents Need Metered Reasoning
Agent quality will depend less on constant reasoning and more on deciding when extra reasoning is worth the cost.
Insurance AI Needs Reasoning-Grounded Retrieval
In insurance, high-risk errors usually come from weak grounding rather than weak language generation.
The Real Bottleneck in Gulf Arabic RAG
Arabic RAG failures frequently come from weak context assembly rather than weak embeddings.
Humility, Verification, and LLM Behavior
What looks like model humility is often an architecture that forces the system to pay for evidence.
Confidence Comes from Skipping the Evidence Loop
LLM confidence is less about model size and more about whether the system is forced to verify.
Why RAG Fails When Retrieval Is Not Evaluated
Many RAG failures are diagnosed as model failures even when the real error happened before generation.
Saudi Enterprises Are Wasting RAG Budget
The expensive part of RAG is not always the model. It is the absence of retrieval discipline.
Why Serious AI Teams Still Need Python
No-code orchestration can be useful at the boundary of a system, but it should not become the center of an AI engineering practice.
LLMs for Legal Compliance Need Memory
Compliance systems need institutional memory: how similar cases were interpreted, escalated, and resolved.
Arabic LLMs Need Evaluation That Understands Context
Arabic model quality should be judged by contextual behavior, not by fluency alone.
Arabic LLM Deployment Needs Semantic Evaluation
Arabic evaluation cannot stop at surface overlap; it must test whether meaning survives dialect, formality, and domain language.
Self-Evolving Agents Raise the Reliability Bar
Adaptive agents require stronger observability because the system that runs today may not behave exactly like the one tested yesterday.