AI Performance Engineering
RepositoriesA curated roadmap from GPU fundamentals and kernel optimization to production inference ~ Still progressing through the resources here, some really great stuff.
A collection of items I've found to be valuable across the internet.
14 bookmarks saved
A curated roadmap from GPU fundamentals and kernel optimization to production inference ~ Still progressing through the resources here, some really great stuff.
An array framework for building and running machine learning workloads on Apple silicon ~ Been messing around with this on my Mac (ported models coming soon).
A walkthrough of using AI as a personalized tutor, with probing, lesson planning, and quizzes ~ It's my inspiration for a learning-focused harness I'm building.
Andrew Ng’s Stanford CS229 lecture series introducing the foundations of machine learning ~ The "Gold Standard" for ML courses.
An animated introduction to LLMs, pretraining, chatbots, and transformers ~ A quick visual grounding for what happens inside an LLM.
A visual walkthrough of the transformer architecture and attention mechanism behind modern language models ~ Awesome visualization of the integral architecture of an LLM.
A case for building deep craft and range while using AI without outsourcing your judgment ~ How to maximize your future in a world that's rapidly changing.
A tour of how language models evolved in the way they store, update, and retrieve information ~ Great breakdown of the evolution of LLMs.
A step-by-step worklog on tiling, memory access, and other CUDA matrix multiplication optimizations ~ One of the first things I read when I started learning CUDA.
A hands-on primer covering tensors, autograd, neural networks, training loops, and GPU use ~ Solid intro to PyTorch fundamentals.
A walkthrough of paged attention, continuous batching, caching, speculative decoding, and distributed LLM serving ~ vLLM deep dive.
An IO-aware attention algorithm that improves GPU utilization through better block and warp work partitioning ~ Significant improvement on the FlashAttention algorithm.
A quantization approach for shrinking LLM KV caches and vector-search embeddings with little accuracy loss ~ Revolutionary quantization method with zero accuracy loss.
Fine-tune language models with low-rank weight updates, reducing trainable parameters and memory without adding inference latency ~ incredibly efficient fine tuning.