I design, evaluate, and deploy AI systems that move from research hypotheses
to reliable, secure, observable, and cost-aware production software.
I am an AI Researcher and Senior AI Engineer working at the intersection of scientific experimentation and production software engineering. My focus is building AI systems that are grounded in evidence, measurable under realistic evaluation, and maintainable after deployment.
My work spans LLM platforms, retrieval-augmented generation, agentic workflows, Arabic and multilingual AI, computer vision, OCR, speech systems, model optimization, secure backend services, and cloud delivery.
| Research Engineering Paper-to-code, reproducible experiments, benchmarking, ablations, claim verification, and evaluation design. |
LLM Systems RAG, GraphRAG, agents, tool use, structured outputs, reranking, guardrails, and human review. |
Production AI Secure APIs, inference services, observability, CI/CD, latency and cost optimization, and cloud deployment. |
Arabic-First AI Arabic NLP, dialect modeling, speech, OCR, RTL-aware interfaces, and multilingual assistants. |
- Start from the problem, data, and measurable success criteriaβnot from a fashionable model.
- Separate prototypes from production systems through explicit reliability, security, latency, and cost requirements.
- Evaluate retrieval quality, hallucinations, edge cases, and failure modes before optimizing headline metrics.
- Use human review when uncertainty, operational risk, or domain impact is high.
- Build observable systems with reproducible experiments, versioned artifacts, and clear rollback paths.
Project-level outcomes reported from the systems described below.
| Outcome | System |
|---|---|
| 60β80% of customer inquiries automated | NexusAI customer experience platform |
| 12% F1 improvement over a Whisper-only baseline | Arabic dialect-aware voice assistant |
| 59% smaller model and 2.3Γ faster CPU inference | Arabic handwriting recognition |
| Approximately 92% character-level accuracy on 50,000 samples | Handwritten prescription recognition |
| 35% reduction in prescription interpretation errors | OCR and verification pipeline |
Research intelligence platform for paper discovery, forensic analysis, claim verification, and paper-to-code generation.
- Built a GraphRAG-based workflow for parsing and analyzing technical papers across text, tables, figures, and OCR-extracted content.
- Developed a Scientific Code Forge that converts paper sections into modular Python and PyTorch project structures.
- Added claim-to-evidence checks that compare extracted claims with reported metrics, tables, and experimental results.
- Designed scientific auditing workflows using adversarial review, hype scoring, and consensus reporting.
- Mapped paper concepts to relevant GitHub implementations for implementation-oriented research.
Stack: Python, Asyncio, Streamlit, ChromaDB, GraphRAG, Docling, PyMuPDF, OCR, OpenAI, Anthropic, Gemini, Ollama
Repository: AI-Research-Intelligence-Agent
Enterprise support automation combining intent classification, RAG, multilingual interaction, and live operations.
- Automated 60β80% of customer inquiries through intent classification and retrieval-grounded response generation.
- Delivered English and Arabic interaction with automatic RTL handling.
- Built workflows for agent takeover, case management, orders, staff administration, and operational monitoring.
- Implemented JWT authentication with HTTP-only cookies and role-based access control.
- Added analytics for latency, error rates, intent distribution, traffic, and operational quality.
Stack: Next.js, React, TypeScript, Node.js, Cloud Run, Vertex AI Gemini, Firestore, BigQuery, Docker
Demo: NexusAI Β· Repository: nexus-ai
Low-latency multilingual voice assistant with Arabic dialect classification.
- Combined Whisper v3, a fine-tuned MARBERTv2 classifier, and downstream LLM reasoning in a modular inference pipeline.
- Targeted sub-500 ms response time for short utterances.
- Improved dialect-recognition F1 by 12% over a Whisper-only baseline.
- Built real-time communication using Flask, React, REST APIs, and WebSockets.
- Orchestrated retries, queues, events, and human review using n8n and Elsa Workflows.
Stack: Whisper, MARBERTv2, Flask, React, WebSockets, n8n, Elsa Workflows
Repository: AIVoiceAsisstent
Deployment-oriented Arabic character recognition with model compression and CPU optimization.
- Trained a CNN on 120,000 handwritten samples across 28 classes.
- Achieved 85% top-1 accuracy using preprocessing, augmentation, and dropout regularization.
- Applied quantization-aware training and ONNX conversion.
- Reduced model size by 59% and improved CPU inference speed by 2.3Γ.
- Exposed the model through a Flask API for public use.
Stack: TensorFlow, Flask, ONNX, PythonAnywhere
Demo: LearnWithUs Β· Repository: LearnWithUs
More applied AI systems
- Designed a multitask pipeline combining GANs, CRNN, and CTC loss for mixed Arabic and Latin handwriting.
- Processed dosage units, physician signatures, and hospital seals.
- Added hospital-seal verification and dosage standardization.
- Reported approximately 92% character-level accuracy on a 50,000-sample dataset and a 35% reduction in interpretation errors.
Stack: GANs, CRNN, CTC Loss, TensorFlow, OCR, Multitask Learning
- Built a retrieval-grounded natural-language-to-SQL assistant for MS SQL Server.
- Added schema validation to reduce hallucinated tables and columns.
- Enforced safety constraints against unsafe or overly broad queries.
Stack: Python, Streamlit, MS SQL Server, Vanna, ChromaDB
- Scientific AI agents, paper analysis, and paper-to-code systems
- GraphRAG, long-context retrieval, reranking, and knowledge-grounded reasoning
- LLM evaluation, hallucination detection, claim verification, and reliability engineering
- Arabic NLP, dialect-aware intelligence, speech systems, and multilingual assistants
- OCR, document intelligence, and multimodal verification pipelines
- Efficient inference, quantization, ONNX, model serving, and cost-aware deployment
Methods: LLMs, RAG, GraphRAG, embeddings, reranking, structured outputs, prompt evaluation, guardrails, OCR, speech AI, computer vision, fine-tuning, PEFT/LoRA, RLHF-style workflows, quantization, and ONNX.
Engineering: ASP.NET Core, Entity Framework, REST APIs, WebSockets, React, FastAPI, Flask, Docker, IIS, CI/CD, OpenTelemetry, Prometheus, Grafana, MLflow, and Weights & Biases.
| Period | Role | Organization | Focus |
|---|---|---|---|
| Dec 2024 β Present | Senior AI Engineer | NVSSoft | AI services, ASP.NET Core, React, CI/CD, IIS, n8n, Elsa Workflows |
| Dec 2025 β Present | AI Training & Model Evaluation Specialist | micro1 | LLM evaluation, rubric design, failure analysis, RLHF/RLAIF workflows |
| Dec 2024 β Dec 2025 | AI Engineer | Reality AI Lab | Generative-image training pipelines, augmentation, evaluation, scalability |
| Sep 2024 β Feb 2025 | Data Scientist Intern | Darrebni | Predictive modeling, SQL analytics, Tableau, feature engineering |
| Sep 2022 β Dec 2024 | AI Engineer | Freelancer.com | NLP, Arabic handwriting recognition, optimization, agentic workflows |
Education and certifications
- M.Sc. Artificial Intelligence, University of Hull β Sep 2025βPresent
Focus: computer vision, deep learning, and LLM applications. - B.Sc. Information Technology Engineering, University of Kalamoon β Aug 2018βFeb 2024
Graduated with distinction; ranked 2nd in class.
| Certification | Issuer | Date |
|---|---|---|
| Develop AI-Powered Prototypes in Google AI Studio | Feb 2026 | |
| AI Software Engineer | micro1 | Oct 2025 |
| Generative AI with Large Language Models | Coursera / AWS | Oct 2024 |
| Introduction to Retrieval Augmented Generation | Duke University / Coursera | Dec 2024 |
| Intermediate Machine Learning | Kaggle | Feb 2025 |
| Feature Engineering | Kaggle | Nov 2024 |
| AWS EMEA Innovate: Migrate. Modernize. Build. | AWS | Oct 2024 |
| Azure DevOps: Intro to CI/CD | United Latino Students Association | Mar 2024 |
Problem framing
β hypothesis and success criteria
β data and evidence audit
β reproducible prototype
β evaluation and failure analysis
β security, latency, cost, and reliability hardening
β monitored deployment
β evidence-driven iteration
Research-backed AI. Production-grade engineering. Systems designed to be evaluated, operated, and trusted.
LinkedIn Β· Email Β· Portfolio Β· Repositories

