Building safety-critical AI/LLM systems at frontier scale.
LLM guardrails across billions of daily interactions Β Β·Β 30K+ TPS distributed systems Β Β·Β MCP Β Β·Β Agent frameworks
AI that is safe, beneficial, and aligned with human values.
Staff-level engineer with 13 years building safety-critical AI/LLM systems at scale. Deep expertise in Responsible AI, backend, and distributed systems β with a focus on Trust & Safety platforms that deliver LLM guardrails across billions of daily interactions.
- π‘οΈ I architect high-throughput distributed systems (30K+ TPS) and embed Safety, Security, Privacy, and Policy directly into production AI products.
- π€ I lead cross-functional initiatives across 20+ teams, taking ambiguous, zero-to-one problems from PRFAQ to production.
- π§ I operate fluidly across macro architecture and micro detail β and I care deeply about building AI that is safe, beneficial, and aligned with human values.
"Changing lives for the better β through technology."
π¦ Online System Safety β Meta AI Muse Spark launch Led end-to-end online safety detection that gated the launch of Muse Spark. Coordinated many cross-functional teams (LLM Trust, Child Safety, Media Trust) to identify and mitigate every risk vector β resolved multiple high-severity issues and closed all CSAM image/video entry points with zero project delay, ramping fast on a novel architecture and codebase. Held refusal rates and latency within target so guardrails never degraded UX.
π¦ Training-Data Filtering Pipeline Built β ground up β a pre-training data-sanitization system for research scientists. Filters CSAM (image, novel video, MMS-bank), foreign-state narrative bias, and other policy violations. Operates at scale within tight SLAs, with a low false-drop rate and fail-close semantics, running many parallel pipelines across large datasets β zero delays to model launches. Coordinated multiple teams and research stakeholders and shipped a companion E2E testing framework.
π¦ MCP Server Platform for Model Evals Designed and deployed a scalable MCP server platform on AWS β an extensible tool-augmented LLM framework where new tools plug in to evaluate model tool-use against safety risks. Established a repeatable, automated safety gate for model releases.
Highlights spanning 13 years across Meta, Amazon, InMobi, and SuccessFactors.
| Scale | Reliability | Efficiency | Leadership |
|---|---|---|---|
| 30K+ TPS safety platforms | 2.6B events/day @ <200 ms | ~50% infra cost reduction | 20+ teams aligned (PRFAQ β prod) |
| Billions of daily interactions | <0.5% policy-violating content | $400K/yr saved via ML automation | Bar Raiser, 350+ interviews |
| Large-scale data filtering | Low false-drop, fail-close | 10x ingestion throughput | 5 interns β FTE engineers |
Alexa AI Β· Privacy & Customer Trust β Senior SDE
- Single-threaded owner of Alexa's end-to-end LLM content-moderation stack across all LLM experiences. Built guardrails β prompt/output overriding hotfixes + pause-and-play full-context detection β scaling to 30K TPS and driving policy-violating content to <0.5% of traffic. Led 5 cross-functional teams.
- Alexa Trust Monitoring System β envisioned and championed a company-wide platform from zero (authored PRFAQ, secured funding, aligned 20+ teams). Fault-tolerant architecture at 30K TPS / 2.6B events/day under 200 ms, established as the authority for Alexa's go/no-go launch decisions.
eCommerce Catalog β SDE II
- Global Item Processing re-architecture β multi-year overhaul of a legacy distributed catalog β 24K TPS @ <100 ms, 50% less artificial traffic, 2x infra cost savings.
- Automated data-quality + ML pipeline β human-in-the-loop (35-person MTurk team) β automated SageMaker training/deploy β 10bps YoY quality gain, $400K/yr saved.
InMobi β Senior SWE Β· real-time ad-serving catalog ingestion (Spring Boot, Elasticsearch): 40 merchants, 10M+ products/day, 10x throughput; Hive audience segmentation 4 days β 2 hours. SuccessFactors (SAP) Β· binary expression-tree permission resolver (ANTLR) for fine-grained admin control β recognized by a Principal Engineer.
Timeline
Overall: 13 years of building safe systems at scale
2013 : SuccessFactors (SAP) β ANTLR permission resolver Β· Upgrade Center
2015 : InMobi β Senior SWE Β· real-time ad catalog Β· 10x ingestion
2017 : Amazon eCommerce Catalog β SDE II Β· 24K TPS re-architecture
2022 : Amazon Alexa β Trust Monitoring Β· 2.6B events/day
2023 : Amazon Alexa β LLM content-moderation owner Β· 30K TPS Β· Bar Raiser
2025 : Meta Superintelligence Labs β Online Safety Β· MCP Β· model evals
- π― Amazon Bar Raiser β 350+ hiring interviews, owning the bar-raising function across engineering orgs.
- π€ GuruConnect β company-wide hackathon winner β AWS Bedrock AI assistant for on-call management, SOP retrieval, and DB querying.
- π° AWS Timestream blog author β co-authored the official post introducing customer-defined partition keys (first production implementation of the feature).
- π Amazon Security Certifier β security-certified 3 production systems/year for customer-trust, privacy, and policy compliance.
- π± Mentorship β grew 5 interns into full-time engineers; ongoing coaching on career growth, system design, and engineering excellence.
π Full rΓ©sumΓ©, detailed experience & recommendations live on my LinkedIn.
π Bellevue, WA Β· π B.E. Computer Science, SJCE Mysuru β 9.21/10 Β· π‘οΈ Safety Β· Security Β· Privacy Β· Policy
