Large language models (LLMs) are the reasoning engines of modern AI. Today, a major inflection point has arrived: as the world races to deploy AI at scale, model inference has moved to the center of the stack. Welcome to the inference era.
Without proper optimization, however, LLMs can be expensive and slow to serve. Hands-On LLM Serving and Optimization is a comprehensive guide to the complexities of deploying and optimizing LLMs at scale.
In this hands-on, engineering-focused book, authors Chi Wang and Peiheng Hu combine practical examples, code, and strategies for building robust, performant, and cost-efficient AI token factories. Whether youâre building the LLM inference infrastructure or the applications that consume it, a deep understanding of LLM serving will make you a more effective, future-ready engineer as AI transforms how we work and build.
Learn the foundations of model serving with core concepts, design paradigms, and industry best practices
Understand the common challenges of hosting LLMs at scale
Balance latency and throughput to meet the demands of AI applications and business requirements
Host LLMs cost-effectively with practical, code-backed techniques
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# Hands-On LLM Serving and Optimization: Hosting LLMs at Scale
## 【One-Line Pitch】
A practical, engineering-focused guide to deploying, serving, and optimizing large language models in production—essential reading for backend engineers, ML infrastructure teams, and technical leads who need to build cost-efficient, high-performance LLM inference systems.
## 【Book Arc】
- **Opening (~0%–9%)**: Establishes the "inference era" thesis—LLM serving has become the critical bottleneck in AI deployment. Introduces the book's scope (not a general ML primer, not a survey of every framework) and defines the target audience: engineers comfortable with Python and basic deep-learning concepts who want production-grade serving knowledge.
- **Early (~9%–25%)**: Covers model serving fundamentals—the critical differences between training and serving frameworks (training optimizes for backpropagation and large batches; serving optimizes for low-latency forward passes). Introduces serving paradigms: on-device, on-premises, and on-cloud deployment scenarios, plus the single-model service pattern and multi-model serving containers with LRU-based model caching.
- **Early (~25%–34%)**: Transitions into LLM-specific serving concepts. Walks through the Transformer architecture from a serving perspective—decoder-only models, autoregressive generation, and the attention mechanism—with hands-on visualization examples using BertViz and the Qwen model.
- **Middle (~34%–44%)**: Deep dive into the mechanics of LLM inference. Explains KV cache as the key optimization that shifts computation from full-sequence recomputation to incremental, cache-augmented workflows. Introduces serving frameworks like vLLM and SGLang as purpose-built systems that handle request scheduling, batching, and token streaming.
- **Middle (~44%–53%)**: Moves from theory to implementation. Demonstrates building a model serving system from scratch using Python multiprocessing—a ModelWorker that monitors task queues, a web API layer, and batching logic that trades individual request latency for significantly improved overall throughput. Concludes with continuous batching, which can improve throughput by up to 23×.
## 【Key Takeaways】
- **Serving frameworks are fundamentally different from training frameworks** (Early): Training optimizes for gradient updates and large batches; serving optimizes for low-latency forward passes on single or small batches. Using training frameworks like PyTorch or Hugging Face Transformers for serving is inefficient—purpose-built tools like vLLM, SGLang, and NVIDIA Triton are designed for production inference.
- **Deployment scenarios shape serving architecture** (Early): Models can be served on-device (edge inference like warehouse robots), on-premises (data-sensitive workloads like document search), or on-cloud. Each scenario has different latency, security, and infrastructure constraints that dictate the serving design.
- **KV cache is the single most important LLM serving optimization** (Middle): By caching attention keys and values from previously generated tokens, the model avoids recomputing attention for the entire sequence during each decoding step. This trades increased memory usage for dramatic compute savings and faster token generation.
- **Serving frameworks provide essential production capabilities beyond raw inference** (Middle): vLLM and similar frameworks handle KV-cache reuse, request scheduling, multi-user concurrency, token streaming, and cancellation/interruption handling—all critical for real-world applications but absent from basic inference implementations.
- **Batching is the primary lever for throughput** (Middle): Static batching processes fixed groups of requests, but continuous batching—where new requests are dynamically added as others complete—can improve LLM inference throughput by up to 23× while significantly reducing p50 latency.
- **Building a serving system from scratch reveals the core design patterns** (Middle): A production serving system requires a web API layer, a workload manager to track requests, and worker processes that pull from task queues. The LLMEngine pattern—collecting prompt IDs, generating results, and cleaning up finished sequences—illustrates how batching optimizes resource utilization across concurrent requests.
## 【Reading Tips】
- **Skim the Transformer architecture review** (Early, ~28%–34%) if you already understand attention mechanisms—the book's value is in the serving perspective, not the ML theory. Focus instead on how attention and KV cache connect to serving decisions.
- **Deep-read the vLLM configuration example** (Early, ~19%): The command-line setup for running gpt-oss-20b on H100 GPUs (with flags like `--gpu-memory-utilization`, `--max-num-seqs`, and `--tensor-parallel-size`) previews the optimization knobs covered in later chapters—understanding these flags early will help you follow the advanced material.
- **Pay close attention to the serving paradigm comparisons** (Early, ~25%): The single-model versus multi-model service patterns, and the role of model cache management with LRU eviction, are foundational design decisions you'll need to make in real deployments.
- **Work through the from-scratch serving implementation** (Middle, ~44%–53%): The Python multiprocessing example with ModelWorker, task queues, and batching logic is the book's most hands-on section. Running this code will cement your understanding of how serving frameworks work under the hood.
- **The book assumes Python proficiency and basic deep-learning knowledge**—if you're new to transformers, pair this with a general LLM introduction before diving in.
## 【Coverage Limits】
The excerpts cover roughly the first half of the book (through ~53%), focusing on fundamentals, serving paradigms, Transformer mechanics, KV cache, and a from-scratch serving implementation. Later chapters on performance measurement, GPU specification analysis, model loading bottlenecks, and advanced optimization strategies (Chapters 5–7) are only previewed in the table of contents and are not covered in this guide.
##
Excerpt 1
125 Core Inference Layer 125 Model Optimization Layer 125 Model Layer 125 Building with an Open Source Stack 126 Implementing Public API 127 Implementing Mod...
SDP] execution, often on CPUs, edge devices, and DeepSpeed). A scale example is the Llama3 405B or inference-specific GPUs (such as NVIDIA model that was tra...
ust a historical curiosity—it offers critical insights into the model design choices, architectural patterns, and execution behaviors, which are fundamental...
t into the result_queue. See the implementation as follows: class ModelWorker: @staticmethod def run(model_name: str, task_queue: mp.Queue, result_queue: mp....
ed to train 1,000 models for 1,000 customers. Since not all models will be used at the same time, you can limit the system to host at most 200 models concurr...
ample of hosting a PyTorch model with TorchServe. The first step is to select a prebuilt DLC that matches your framework and Python version. You can use the...
apter, you will have the intuition you need to navigate the world of LLM optimization with confidence. The next chapters will build upon this foundation by g...
o 512, the arithmetic intensity is 21 and 170, respectively. Recall the calculation of the roofline model in Figure 5-14: the crossover point for L40S is 419...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
Hands-On LLM Serving and Optimization Hosting LLMs at Scale (Chi Wang, Peiheng Hu)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
Hands-On LLM Serving and Optimization Hosting LLMs at Scale (Chi Wang, Peiheng Hu)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment