I got into data engineering by keeping production pipelines alive at a manufacturing company — 57+ Airflow DAGs feeding metrics that business teams opened every morning. The AI side came from working on RAG systems, where the real problem isn't getting a model to answer — it's getting it to answer with evidence you can check.
- AI Engineer Intern @ Sonus Software Solutions (2025–26) — took an internal RAG Q&A agent from FAISS prototype to production on Pinecone, over a 10,000+ document corpus
- Associate Data Engineer @ NSR Industries (2022–24) — production Airflow + dbt pipelines, 30+ pipeline failures diagnosed across two on-call rotations
- MS in Business Analytics & AI, The University of Texas at Dallas (2026)
- Databricks Certified Data Engineer Associate
A Random Forest classifier flags attacks across 2.8M+ real network flows at 94.1% F1 — and a RAG layer grounds every detection in MITRE ATT&CK evidence, so each flagged threat comes with a justification instead of a bare confidence score.
PySpark BigQuery dbt Dagster scikit-learn MLflow Qdrant Claude API FastAPI Docker GCP Terraform
Actively expanding: real-time streaming inference over Kafka, a RAG evaluation harness measuring retrieval accuracy and citation faithfulness, and a GCP migration provisioned with Terraform.
Scores incoming transactions with LightGBM, then decides approve/decline on expected cost, not a fixed threshold — because declining a loyal customer and approving a $1,200 fraud are not the same mistake. Kafka streaming ingestion into an S3/Iceberg lakehouse, modeled in Snowflake + dbt, serving decisions at p50 ~28ms.
Kafka AWS Apache Iceberg Snowflake dbt Airflow LightGBM DynamoDB FastAPI Terraform GitHub Actions React
Actively expanding: Databricks + Unity Catalog model registry with MLflow experiment tracking, and live PSI-based drift detection wired into the API's health endpoint.
AI & ML
Data Engineering
Cloud & Infra
Languages & Backend
Visualization
Clarity is powerful. Efficiency is underrated.