GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization for High-Throughput AI Production Systems (AI Infrastructure, Hardware & Compiler Engineering Series)
I will be using this book for:

GPU Kernel Engineering for LLM Inference: CUDA, Triton, and Flash Attention Optimization for High-Throughput AI Production Systems (AI Infrastructure, Hardware & Compiler Engineering Series)

by Chatvariety Team

Technology Programming Computer Science artificial intelligence
1 Star 2 Star 3 Star 4 Star 5 Star
0.0 out of 5 stars (0 ratings)

Explore GPU kernel engineering for large language model inference with coverage of CUDA, Triton, and Flash Attention optimization. This technical guide centers on AI infrastructure, hardware, compiler engineering, and the performance considerations behind high-throughput production systems.

About This Book

GPU Kernel Engineering for LLM Inference focuses on optimizing large language model inference for high-throughput AI production systems.

The book covers CUDA and Triton, two technologies associated with GPU programming and kernel development.

It also addresses Flash Attention optimization as part of the broader work of improving inference performance.

Designed around AI infrastructure, hardware, and compiler engineering, this title presents a technical perspective on production-oriented systems.

Reviews

No reviews yet. Be the first to review this book!


Write a Review
I will be using this book for: