GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization for High-Throughput AI Production Systems (AI Infrastructure, Hardware & Compiler Engineering Series)
Explore GPU kernel engineering for large language model inference with coverage of CUDA, Triton, and Flash Attention optimization. This technical guide centers on AI infrastructure, hardware, compiler engineering, and the performance considerations behind high-throughput production systems.
About This Book
GPU Kernel Engineering for LLM Inference focuses on optimizing large language model inference for high-throughput AI production systems.
The book covers CUDA and Triton, two technologies associated with GPU programming and kernel development.
It also addresses Flash Attention optimization as part of the broader work of improving inference performance.
Designed around AI infrastructure, hardware, and compiler engineering, this title presents a technical perspective on production-oriented systems.
Reviews
No reviews yet. Be the first to review this book!