In today's era of ever-growing generative models, AI Systems Performance Engineering equips professionals with actionable strategies to co-optimize hardware, software, and algorithms for high-performance and cost-effective AI systems. Authored by Chris Fregly, a performance-focused engineering and product leader, this comprehensive resource transforms complex systems into streamlined, high-impact AI solutions. Whether you're an engineer, researcher, or developer, this book offers a holistic roadmap for building resilient, scalable, and cost-effective AI systems that excel in both training and inference.
AI Reading Assistant
Whole-book reading guide from stratified index samples; jump to passages in the text
Tip the Site
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat Pay
Alipay
Open WeChat or Alipay and scan. No login required.
AI guide
# AI Systems Performance Engineering
## 【One-Line Pitch】
A practical, full-stack guide to co-optimizing hardware, software, and algorithms for AI systems—covering everything from NVIDIA's latest GPU architectures to Kubernetes networking tweaks—for engineers, researchers, and developers who want to squeeze maximum performance and cost-efficiency out of training and inference workloads.
## 【Book Arc】
- **Opening (~0%–10%)**: Establishes the "why" of AI systems performance engineering through the DeepSeek case study—a $6M training run that rivaled models costing orders of magnitude more—and introduces core principles like profile-driven optimization, holistic system views, and striving for order-of-magnitude impact rather than incremental gains.
- **Early (~10%–23%)**: Dives deep into NVIDIA's hardware roadmap, explaining the Grace-Blackwell superchip architecture with its unified memory (up to 864 GB per superchip), the NVL72 rack-scale system with 72 GPUs interconnected via NVLink, and the progression toward GB300 with Blackwell Ultra GPUs pushing into exascale territory.
- **Early (~23%–32%)**: Transitions from hardware to software optimization, covering CPU and OS tuning techniques—NUMA awareness, CPU pinning with taskset, memory pinning for faster DMA transfers, huge pages, and disabling C-states to eliminate latency-inducing OS interference.
- **Middle (~32%–42%)**: Explores GPU sharing and partitioning strategies including Multi-Process Service (MPS) for overlapping work from multiple processes, Multi-Instance GPU (MIG) for hardware-level partitioning, and Kubernetes time-slicing—plus container optimization and host networking for multi-node GPU workloads.
- **Middle (~42%–48%)**: Focuses on data flow optimization—prefetching with PyTorch's DataLoader, batching I/O operations, overlapping communication with computation using CUDA streams, and network tuning for Ethernet and InfiniBand interconnects.
- **Middle (~48%–end)**: Covers inference-specific optimizations including KV-cache management for Transformer models, the distinction between compute-bound prefill and memory-bound decode stages, and advanced techniques like NVMe-based KV-cache offloading for distributed inference clusters.
## 【Key Takeaways】
- **Performance engineering is a multiplier, not a luxury** (Opening): DeepSeek's $6M training run versus billion-dollar competitors proves that clever system optimizations can upend AI economics—at scale, every bit of performance translates to millions saved. (Early)
- **Unified memory architectures are game-changers** (Early): The Grace-Blackwell superchip's cache-coherent NVLink-C2C interconnect lets CPU and GPUs share nearly a terabyte of memory as one pool, eliminating explicit data copies and enabling models that previously couldn't fit on a single GPU. (Early)
- **Lower precision formats double throughput at each step** (Early): FP8 doubles FP16 throughput, and Blackwell's new FP4 doubles FP8 again—one Blackwell GPU hits ~9 PFLOPS with FP4, roughly 4× its FP16 rate, making mixed-precision training a core performance lever. (Early)
- **CPU tuning is as important as GPU selection** (Early): Pinning processes to cores with taskset, aligning memory allocations within NUMA nodes, using pinned memory (2–3× faster GPU transfers), and disabling C-states can eliminate the "bubbles" where GPUs sit idle waiting for data. (Early)
- **GPU sharing requires matching strategy to workload** (Middle): MPS overlaps execution from multiple processes for throughput, MIG provides hardware-isolated partitions for multi-tenant environments, and Kubernetes time-slicing offers simple isolation at the cost of idle time—choose based on whether you prioritize isolation or utilization. (Middle)
- **Host networking eliminates container overhead for GPU clusters** (Middle): Setting hostNetwork: true in Kubernetes lets containers access InfiniBand directly without NAT or overlay translation—critical for MPI jobs and performance-sensitive multi-node workloads. (Middle)
- **Overlapping communication and computation keeps GPUs fed** (Middle): Using CUDA streams to run all-reduce gradient aggregation alongside matrix multiplications, plus prefetching data batches ahead of time, reduces idle GPU time and improves overall throughput. (Middle)
- **Inference has two distinct stages requiring different optimizations** (Middle): Prefill is compute-bound (building KV-cache from prompts), while decode is memory-throughput-bound (gathering weights for token generation)—advanced systems like NVIDIA Dynamo and vLLM run these stages on separate GPUs, and KV-cache offloading to NVMe handles idle sessions. (Middle)
## 【Reading Tips】
- **Skim the hardware roadmap sections** (~10%–23%) if you're not selecting GPU infrastructure—the key takeaway is understanding how unified memory and lower precision create optimization opportunities, not memorizing specific specs.
- **Deep-read the CPU/OS tuning chapter** (~29%–32%)—the practical techniques (taskset, numactl, pin_memory=True, ulimit settings) are immediately applicable to any GPU server you'll work with.
- **Pay special attention to the GPU sharing section** (~32%–39%)—MPS vs. MIG vs. time-slicing is a decision you'll face in any multi-tenant environment, and the trade-offs are clearly explained.
- **The inference section** (~48%+) is essential if you're deploying models for production—understanding prefill vs. decode stages and KV-cache management is critical for building efficient inference servers.
- **Take away the profiling mindset**: the book repeatedly emphasizes measuring everything, trusting data over assumptions, and using profilers to identify true bottlenecks before applying targeted optimizations.
## 【Coverage Limits】
Excerpts do not cover specific code examples for implementing the discussed techniques, nor do they detail the mathematical foundations of Transformer architectures or provide step-by-step benchmarking methodologies.
##
Page 16
eepSeek claims that their DeepSeek-R1 model was trained for only $6 million in compute – an order of magnitude lower than models like GPT-4 and Gemini – whil...
l GPUs - and 36 Grace CPUs - all interconnected via NVLink. The GB200 NVL72 is built as 18 compute nodes (each 1U in size), where each node contains two GB20...
g that might introduce unpredictable latency such as excess context switching, frequency scaling, and memory-to-disk swapping. The result should be that your...
based network, you can use the Linux networking /etc/sysctl.conf parameters: net.core.rmem_max , net.core.wmem_max to set max buffer sizes, and net.ipv4.tcp_...
onal headroom to scale the number of end users even higher. With the same number of GPUs, Dynamo’s dynamic batching, GPU load balancing, and disaggregated in...
educe the learning rate” as shown in Figure 6-3. Figure 6-4. Dynamic graph execution based on input data (Source: https://developer.nvidia.com/blog/enabling-...
ectivity for multi-GPU workloads. NVLink 5 provides up to 1.8 TB/s GPU-to-GPU Ensure your GPU servers run a recent, stable Linux kernel configured for high-p...
over FP16. On newest GPUs, also explore and evaluate FP8 or INT4 for certain models to further boost throughput for inference. Fused Activation + Scaling. Wh...
Support this siteYour recognition and a small knowledge-service contribution help keep this technical work open source.
Scan the WeChat Pay or Alipay code below. Logged-in and guest visitors can both tip.
WeChat PayAlipay
Open WeChat or Alipay and scan. No login required.
Add Tag
Enter tag name (max 50 characters)
Share E-Book
AI Systems Performance Engineering (for Raymond Rhine) (Chris Fregly)(Z-Library)
Scan QR code with your phone to access
Copy the link or scan the QR code to access this e-book on your phone
Share E-Book via Email
Please enter email address
Donation Statistics
¥.00
Total Donations
0
Donation Count
AI Systems Performance Engineering (for Raymond Rhine) (Chris Fregly)(Z-Library)
Find Your Favorite Books
Only registered users can comment after logging in. Comments need to be reviewed by administrators before being displayed
Loading comments...
Reply to Comment
Edit Comment