Evaluation Before Automation
I start with what must be proven: evidence recall, decision correctness, exception behavior, citation faithfulness, audit trail quality, and conditions for abstention.
Professional record
Most teams can access the same models. My work is deciding what can be trusted, what must be verified, and how AI becomes an auditable system inside regulated Arabic-English workflows.
Non-commodity edge
The durable advantage is not tool access. It is the ability to turn uncertain model behavior into evidence, controls, review points, and operating routines that a real institution can defend.
I start with what must be proven: evidence recall, decision correctness, exception behavior, citation faithfulness, audit trail quality, and conditions for abstention.
Insurance and compliance work require authority boundaries, human review, policy interpretation, exception handling, and a record of why a decision was made.
Bilingual systems fail in specific ways: translated clauses drift, local terminology matters, dialect changes intent, and document hierarchy can decide the answer.
Doctoral research and first-author publications shaped a habit of controlled comparison, reproducibility, and skepticism toward unsupported model claims.
Evidence record
My background combines a PhD in Electrical Engineering, first-author research in computer vision and optimization, and delivery experience across insurance, compliance-oriented automation, and bilingual Arabic-English enterprise environments.
Electrical Engineering; machine learning, optimization, and computer vision
first-author peer-reviewed AI publications
regulated enterprise AI delivery in the Saudi market
Arabic-English evaluation, retrieval, and workflow context
Global Insurance Innovation Forum Saudi Arabia 2025
A keynote and executive panel on why insurance transformation succeeds only when data quality, change management, legacy integration, and AI controls are treated as one operating problem.
Read field noteSelected systems
The emphasis is not demo fluency. The emphasis is whether a system remains legible when policies conflict, evidence is incomplete, and decisions carry operational consequence.
Work on underwriting, claims, fraud indicators, audit trails, and human review. The technical challenge is inseparable from policy interpretation, authority, exception handling, and operational adoption.
Architecture for systems where retrieval, tool calls, memory, policy constraints, model evaluation, and auditability are treated as one control surface.
Evaluation for systems that must handle Arabic and English documents, translated policy language, local terminology, dialectal intent, and evidence hierarchy.
First-author research on ensemble optimization and histology image classification, with emphasis on reproducible experimentation and model behavior under constraint.
Essays
Essays on RAG, agents, MCP, insurance automation, Arabic-English evaluation, and the control systems around AI.
Broker automation needs a control layer, not another set of screens wired together.
Compute accounting is becoming a governance record, not only an infrastructure metric.
MCP should be understood as an interface discipline: it makes context and tools explicit enough to govern.
MCP is useful only when teams treat context, tools, and session state as governed interfaces.
The hard problem is not calling one tool; it is preserving meaning across a sequence of stateful tool interactions.
MCP turns ad hoc agent integrations into contracts that can be tested, governed, and changed.