AI Engineer working across the stack on RAG, LLM fine-tuning, and real-time speech systems. My work has moved from computer-vision and speech models toward applied GenAI β fine-tuning open LLMs, building retrieval pipelines on vector DBs, and shipping low-latency audio workflows to cloud. Based in San Francisco, CA.
- AI Fund β Technical Builder / AI Engineer (Dec 2025βpresent, Mountain View, CA)
- Ditto AI β AI Engineer (MayβDec 2025, Berkeley, CA)
- JetskiAI β Founding AI/ML Engineer (MarβDec 2025, SF Bay Area)
- SuperIntro β AI Software Engineer (Dec 2024βApr 2025, SF) β Fine-tuned Qwen 2.5 LLM and Stable Diffusion pipelines on Vertex AI with LightRAG; deployed fine-tuned models on GCP with cloud logging and monitoring; integrated via Azure AI Foundry.
- Sizzle β AI Engineer (JanβMar 2025, SF) β Designed a low-latency Whisper + Qwen 2.5 + BERT workflow with Librosa/TorchAudio feature extraction; +15% metadata-tagging precision, +20% acoustic-linguistic alignment.
- Melp App, Inc. β Software Developer (AI/ML) (MayβJun 2025, SF Bay Area)
- Seattle University β Research Assistant (Aug 2024βDec 2025, Seattle, WA) β Under Prof. Pejman Khadivi: fine-tune Transformers and CNNs for NLP and predictive analytics, with emphasis on automation and model deployment.
- Seattle University β Teaching Assistant, Visual Analytics (MarβJun 2024, Seattle, WA)
- SlashRTC β Machine Learning Engineer (Sep 2021βAug 2022) / ML Intern (JunβAug 2021), Mumbai, India β Built Speech-to-Text models with Python and TensorFlow.
|
Problem: Bridge communication for the deaf and hard-of-hearing by translating American Sign Language signs into spoken audio. Approach: Fine-tune an I3D (Inflated 3D ConvNet) on the WLASL dataset for word-level sign recognition, piping predictions into a TTS stage. I3D captures spatiotemporal features across stacked video frames rather than treating frames independently, which suits the motion-heavy nature of signing. Stack: PyTorch Β· I3D Β· WLASL Β· OpenCV |
Problem: Enable live cross-language conversation without the stop-and-wait of batch translation. Approach: A streaming pipeline chains Whisper (ASR) β translation β OpenAI TTS, with audio streamed in and out continuously. It prioritizes low end-to-end latency by keeping the stages pipelined rather than processing each utterance as a discrete block. Stack: Whisper Β· OpenAI TTS Β· Python Β· streaming audio I/O |
Languages
ML / Deep Learning
GenAI / LLM
Speech / Audio Β· Vision
Cloud / Infra
I enjoy giving back to the AI/ML community and am always happy to:
- π Judge hackathons, demo days, and AI/ML competitions
- π€ Give interviews, talks & guest sessions on applied GenAI, RAG, and real-time speech
- π§ Mentor & guide engineers and students breaking into AI/ML
- π‘ Consult & advise on AI/ML product direction and architecture
π« Reach me at rakeshutekar60@gmail.com or on LinkedIn.
- MS, Computer Science (Data Science Specialization) β Seattle University (Sep 2022βAug 2024)
- BTech, Computer Engineering β University of Mumbai (2016β2021)





