Inference at Full Throttle : LLM serving performance with vLLM, quantization, KV cache tuning and speculative decoding
Overview
Master LLM Inference and Scale Your AI Infrastructure
In 2026, inference spend surpassed training spend across the tech industry. The engineers who can maximize tokens per second on H100, H200, and B200 GPU fleets are the most valuable specialists in AI. Inference at Full Throttle turns complex GPU performance engineering into a reproducible, highly practical discipline.
Written by the ChatVariety Team-an elite collective of ML infrastructure engineers and vLLM contributors-this book provides the exact mathematical formulas and production configurations needed to run large language models at extreme scale without breaking the bank.
What You Will Master:- GPU Memory Mathematics: Derive the exact KV cache formula from first principles to budget HBM memory perfectly.
- Advanced Quantization: Deploy FP8, AWQ INT4, GPTQ, and MXFP4 based on real-world throughput and quality trade-offs.
- Serving Stack Optimization: Fine-tune vLLM, SGLang, TensorRT-LLM, and TGI for enterprise workloads.
- Ultra-Fast Decoding: Implement speculative decoding, EAGLE-class self-speculation, and KV cache prefix caching.
- Distributed Scale: Combine tensor, pipeline, and expert parallelism (MoE) across multi-node GPU clusters.
- Production Benchmarking: Avoid common traps by measuring TTFT, TPOT, and tail latency under realistic workloads.
Stop wasting millions on sub-optimal cloud GPU allocations. Learn how to design, benchmark, and operate multi-tenant, high-throughput, and ultra-low-latency LLM serving architectures today.
This item is Non-Returnable
Customers Also Bought
Details
- ISBN-13: 9798192412626
- ISBN-10: 9798192412626
- Publisher: Independently Published
- Publish Date: August 2026
- Dimensions: 9 x 6 x 0.17 inches
- Shipping Weight: 0.27 pounds
- Page Count: 82
Related Categories
