███████╗██╗ ██╗██╗ ███████╗██╗██╗ ██╗███████╗██████╗
╚══███╔╝██║ ██║██║ ██╔════╝██║██║ ██╔╝██╔════╝██╔══██╗
███╔╝ ██║ ██║██║ █████╗ ██║█████╔╝ █████╗ ██████╔╝
███╔╝ ██║ ██║██║ ██╔══╝ ██║██╔═██╗ ██╔══╝ ██╔══██╗
███████╗╚██████╔╝███████╗██║ ██║██║ ██╗███████╗██║ ██║
╚══════╝ ╚═════╝ ╚══════╝╚═╝ ╚═╝╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝
I build things close to the metal — GPU kernels, distributed training infrastructure, and systems software. I care about what happens at the layer most engineers treat as a black box: how memory moves, how schedulers make decisions, how kernels actually run on hardware.
Currently going deep on ML systems, GPU programming, and low-level systems (OS internals, networking, memory). Previously built full-stack apps; now I live in the backend.
A cluster-level job scheduler for heterogeneous GPU environments.
|
Implementing Flash Attention from scratch using Triton GPU kernels, then training GPT-2 on top to validate correctness and benchmark real-world throughput.
|
|
Languages |
ML & GPU |
Infra |
|
Systems |
Previously |
DSA |