I'm a Senior MLE & MLOps Engineer and a Ph.D. researcher from Brazil 🇧🇷, and I've never treated those two as separate careers.
The research side is where I learned to be critical and rigorous about method, evidence and the limits of what a model can actually tell you; the engineering side is where that turns into systems people depend on in production.
Day to day I work across the whole model lifecycle: from arguing about whether a problem is really a classification problem, to keeping the endpoint healthy six months after launch. Most of that happens on Databricks and AWS, with earlier work on GCP and Azure. I do it with a bias toward ML systems that are reproducible, observable and boring to operate. The interesting part should be the modeling and the problem it solves, not the pipeline duct tape.
- 🎓 Academic track: Ph.D. in Organizational Studies, with research spanning computational and quantitative methods (NLP, multivariate statistics) and the social dimensions of technology
- 🔄 End to end: model selection, experiment tracking, registry, deployment, serving, monitoring and the CI/CD that connects it all
- ⚡ Serving tiers: batch, near real time and real time; picking the latency and cost profile the use case actually needs, not the one that sounds impressive
- 🧱 Platform work: repository scaffolding and templates, IaC with Terraform/Terragrunt, developer experience through internal portals (Backstage)
- 🎤 Teaching & speaking: former teacher, and a recurring face at Python community events
- ✍️ Writing: I document my learnings on Medium, from architecture decisions to tooling and certifications
- 💬 Ask me about: Python, PySpark, Databricks, AWS, MLflow, Terraform, Docker, deployment strategies and drift detection
Where I can actually help, stage by stage:
| Stage | What I bring |
|---|---|
| Framing & model choice | Matching the problem to the right family (classification, regression, clustering, forecasting, NLP) and calling out when a simpler baseline would do |
| Experimentation | Cross-validation design that respects time and group leakage, hyperparameter tuning, honest evaluation metrics |
| Registry & versioning | MLflow tracking and registry, model signatures, tagging conventions, promotion flows between environments |
| Deployment | Blue/green, canary and shadow rollouts, rollback paths, and choosing which one the risk profile justifies |
| Serving & sizing | Endpoint design across batch, NRT and RT; sizing compute units, memory and datastores against latency and cost budgets |
| Monitoring | Data and concept drift detection, performance decay, alerting, and deciding what actually warrants a retrain |
| Underneath it all | Repository scaffolding, CI/CD for ML, dependency and environment management, IaC, and service catalogs that make the paved road the easy road |
I work across proprietary and open stacks, and I lean on open tooling whenever it makes the platform easier to audit and reuse. Open source is a direction here, not a badge.
| Languages & OS | Cloud & Data Platform | IaC, CI/CD & DevEx | Data & ML |
|---|---|---|---|
Python |
Databricks |
Terraform |
PySpark |
Linux |
AWS |
Docker |
MLflow |
Bash |
Azure |
GitHub Actions |
Scikit-learn |
Datadog |
Backstage |
Pandas |
- Serving Endpoints no Databricks
- Microbatch, NRT ou RTM no DB
- Terragrunt 101
- Shell para Pythonistas: Episódio II (data handling)
- Tipos de Encoding em ML — Parte II: Encoders Ordinais

