Share E-Book
Scan to open this page

Scan with your phone to open this page

Author: James Brookbank, Leah Hargreaves, Salim Virji

Rating No ratings yet

What happens when DevOps principles start to bend under the weight of scale? At Google and many other large organizations, a new discipline emerged to answer that question: developer platform engineering. This book combines lessons from years of experience with actionable patterns and guidance for teams building or evolving internal platforms. You'll learn what worked and what didn't, when to centralize versus decentralize, and how to apply product thinking to developer experience. Whether you're launching a new initiative or refining an existing one, this book will help you establish the right foundations to scale DevOps sustainably.

AI Reading Assistant

Whole-book reading guide from stratified index samples; jump to passages in the text

AI guide
# Platform Engineering at Google ## 【One-Line Pitch】 A practical guide from Google engineers on building internal developer platforms that scale DevOps sustainably, covering everything from foundational concepts to real-world implementation patterns. Essential reading for platform teams, DevOps practitioners, and engineering leaders at organizations where DevOps principles are straining under scale. ## 【Book Arc】 - **Opening (~0%–10%)**: Introduces platform engineering as the discipline that emerged when DevOps principles broke under organizational scale, with the book's structure spanning Google's internal experience, cloud offerings, and generalizable patterns. - **Early (~10%–23%)**: Explores Google's internal platform evolution through concrete examples like Borg (the predecessor to Kubernetes), Google-Wide Profiling (GWP), and AutoFDO, showing how centralized infrastructure creates leverage through shared optimization. - **Early (~23%–32%)**: Details the "shifting optimizations down" philosophy—moving performance improvements into the platform itself so teams don't rediscover the same lessons, including the SwissMap hash table story and on-by-default optimization strategies. - **Middle (~32%–48%)**: Covers performance observability infrastructure, application productivity metrics (APMs), understanding workload criticality, and the trade-offs of configurability versus standardization in platform design. - **Late (~48%–100%)**: Addresses ecosystem thinking, the dangers of assuming platform immutability, and the socio-technical aspects of platform engineering, including shared fate and responsibility models. ## 【Key Takeaways】 - **Centralized infrastructure creates compounding leverage** (Early): When common libraries and platforms are optimized once, every team benefits without individual effort—as demonstrated by AutoFDO doubling PGO adoption in its first year and TCMalloc optimizations benefiting the entire fleet. - **On-by-default optimization beats opt-in** (Early): Making improvements automatic by default, with rare opt-outs for significant regressions, reduces complexity and sustains high change cadence. Options increase state-space that must be considered with every future change. - **Shift performance left, not just correctness** (Early): Performance regressions are easier and cheaper to fix earlier in the development lifecycle. Techniques include fleetwide telemetry alerts, canary deployments with automated checks, periodic benchmarks, and automated code review tooling. - **Platform-wide migrations can unlock ecosystem dividends** (Early): The SwissMap example shows how centrally-driven refactoring at scale (using large-scale change infrastructure) delivers benefits no individual team could achieve alone, while also making the "right choice" the easy choice for new code. - **Observability infrastructure creates a virtuous cycle** (Middle): Consistent infrastructure (like a common RPC system) makes it possible to add tracing, profiling, and telemetry incrementally, which in turn drives more investment in observability—avoiding the "streetlamp effect" of only optimizing well-understood workloads. - **Application productivity metrics (APMs) connect resources to value** (Middle): Measuring "useful work" rather than raw resource usage reveals true efficiency opportunities, as shown by TCMalloc's "faster than light" optimization that sped up applications beyond what naive allocator improvements could achieve. - **Configurability is a double-edged sword** (Middle): While options seem helpful short-term, they become a tax on velocity long-term. The SwissMap randomization example shows how a small direct cost (randomized iteration order) preserves the ability to make future optimizations. - **Don't assume the platform is immutable** (Middle): Teams often work around platform limitations instead of fixing them once for everyone. Discovering and surfacing these pain points can lead to ecosystem-wide improvements that benefit all participants. ## 【Reading Tips】 - **Skim the early chapters** (~0%–10%) if you're already familiar with platform engineering concepts; the real value starts with the concrete Google examples around Borg, GWP, and AutoFDO. - **Deep-read the "shifting optimizations down" section** (~13%–23%)—this is the intellectual core of the book, showing how platform teams can systematically move best practices into shared infrastructure. - **Pay special attention to the SwissMap case study** (~19%–23%): it's a complete example of the full pattern—identifying a best practice, scaling it across the monorepo, and managing the ecosystem consequences. - **The APM and criticality discussion** (~39%–48%) is dense but crucial for anyone designing platform metrics or efficiency programs; take time to understand the "faster than light" concept. - **Watch for the trade-off frameworks** throughout—the book consistently presents both sides (centralization vs. decentralization, configurability vs. standardization) rather than dogmatic answers. ## 【Coverage Limits】 The excerpts focus heavily on efficiency, performance optimization, and platform infrastructure patterns. The guide does not cover the book's later chapters on security deep dives, AI with platform engineering, or socio-technical platforms in detail, as those sections were not included in the source material. ##
Page 3
m/catalog/errata.csp?isbn=9781098169435 for release details. The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Platform Engineering at Goog...
View in text
Excerpt 2
itecture will affect them and prepare for it. On-by-Default In developing these optimizations, we generally strived for an “on by default” approach. After pr...
View in text
Page 15
ething will work at all. Creating Performance Observability The consistent infrastructure of our ecosystem has allowed us to steadily improve observability,...
View in text
Page 18
ties are effectively a cost of doing business. It’s hard to run a reliable, low-latency service distributed across numerous servers without them. Nonetheless...
View in text
Excerpt 5
This has saved a significant percentage of fleet CPU usage. Similarly, security is treated as an emergent property of the platform; the central “Safe Coding”...
View in text
Excerpt 6
ed systems (a.k.a high shared-fate environments like Type 4 ecosystems) aim for assurance as an explicit design goal. Here, the ecosystem utilizes “safe-by-d...
View in text
Excerpt 7
requires acknowledging that these are inextricably linked. By making these expectations explicit, the stakeholder teams can provide a mutually-agreeable fram...
View in text
Excerpt 8
conceive of value, risk, and effort in software engineering. Whether operating in the highly integrated “walled garden” of a highly integrated ecosystem or t...
View in text
Tags
AI categories
Cloud NativeDevOpsBackend
Publish Year: 2026
Language: English
Pages: 55
File Format: PDF
File Size: 5.0 MB
Text Preview (First 20 pages)
Registered users can read the full content for free

Register as a Gaohf Library member to read the complete e-book online for free and enjoy a better reading experience.

Generating text preview…