Make Programming Simple Lab
Undergraduate Researcher
Researching GPU kernels and compilers to accelerate ML and HPC.
- Developing a torch.compile pass that routes FP64 matmuls between native execution and Ozaki Tensor Core emulation.
- Writing custom PyTorch operators that expose Ozaki kernels and integrate them into the PyTorch backend.
- Building accuracy and cost models to optimize emulation scheme, slice count, and backend kernel strategy per workload.
- Integrating CUDA/Triton fused kernels and benchmarking against cuBLAS FP64 and existing emulation methods.




