Hey, I'm Ammar
ML Systems & Inference Engineer
I build the stuff that actually serves LLMs in production — continuous batching, KV-cache-aware scheduling, quantized kernels, the whole runtime path. From single GPUs to multi-node clusters, I care about the unsexy constraints: latency, cost, utilization, and not melting the hardware.
Currently shipping inference engines on vLLM/SGLang-class stacks, profiling with Nsight, writing C++/Triton/CuteDSL, and occasionally fixing upstream CUDA things (CUTLASS PR #3589).
- Chimera — SGLang fork with a TileLang-first kernel strategy (MLA decode, FP8 GEMM, expert-grouped kernels on Hopper/Blackwell)
- Mercury — Native C++ DAG scheduler + threadpool (GIL-free critical path, pybind11, pure-Python fallback)
If you’re fighting the same problems — kernel dispatch overhead, memory layout, or just trying to keep GPUs busy without lighting money on fire — say hi.
⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣤⡶⠿⠿⠷⣶⣄⠀⠀⠀⠀⠀
⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣰⡿⠁⠀⠀⢀⣀⡀⠙⣷⡀⠀⠀⠀
⠀⠀⠀⡀⠀⠀⠀⠀⠀⢠⣿⠁⠀⠀⠀⠘⠿⠃⠀⢸⣿⣿⣿⣿
⠀⣠⡿⠛⢷⣦⡀⠀⠀⠈⣿⡄⠀⠀⠀⠀⠀⠀⠀⣸⣿⣿⣿⠟
⢰⡿⠁⠀⠀⠙⢿⣦⣤⣤⣼⣿⣄⠀⠀⠀⠀⠀⢴⡟⠛⠋⠁⠀
⣿⠇⠀⠀⠀⠀⠀⠉⠉⠉⠉⠉⠁⠀⠀⠀⠀⠀⠈⣿⡀⠀⠀⠀
⣿⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢹⡇⠀⠀⠀
⣿⡆⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣼⡇⠀⠀⠀
⠸⣷⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢠⡿⠀⠀⠀⠀
⠀⠹⣷⣤⣀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣀⣰⡿⠁⠀⠀⠀⠀
⠀⠀⠀⠉⠙⠛⠿⠶⣶⣶⣶⣶⣶⠶⠿⠟⠛⠉⠀⠀⠀⠀⠀⠀




