botbw here 👋
I'm learning MLSys-related topics, including distributed training frameworks, compilers, and GPU kernels.
I enjoy open sourcing and have been contributing to several projects. Check them out below!
Get to know me through my code!
Making large AI models cheaper, faster and more accessible
A Python library transfers PyTorch tensors between CPU and NVMe
Automated Parallelization System and Infrastructure for Multiple Ecosystems
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
HieraSparse: Hierarchical Semi-Structured KV-Cache Attention on Sparse Tensor Core
Cuda 5