Page
1
(This page has no text content)
Page
2
Site Reliability Engineering SECOND EDITION How Google Runs Production Systems With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write—so you can take advantage of these technologies long before the official release of these titles. Betsy Beyer, Chris Jones, Christof Leng, David Huska, Jennifer Petoff, and Niall Richard Murphy
Page
3
Site Reliability Engineering by Betsy Beyer, Chris Jones, Christof Leng, David Huska, Jennifer Petoff, and Niall Richard Murphy Copyright © 2027 Google LLC and Niall Murphy. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://oreilly.com). For more information, contact our corporate/institutional sales department: 800-998-9938 or corporate@oreilly.com. Acquisitions Editor: Megan Laddusaw Development Editor: Jill Leonard Production Editor: Clare Laylock Interior Designer: David Futato Cover Designer: Karen Montgomery Illustrator: Kate Dullea April 2016: First Edition November 2026: Second Edition Revision History for the Early Release 2026-03-05: First Release See http://oreilly.com/catalog/errata.csp?isbn=9798341607682 for release details.
Page
4
The O’Reilly logo is a registered trademark of O’Reilly Media, Inc. Site Reliability Engineering, the cover image, and related trade dress are trademarks of O’Reilly Media, Inc. The views expressed in this work are those of the authors and do not represent the publisher’s views. While the publisher and the authors have used good faith efforts to ensure that the information and instructions contained in this work are accurate, the publisher and the authors disclaim all responsibility for errors or omissions, including without limitation responsibility for damages resulting from the use of or reliance on this work. Use of the information and instructions contained in this work is at your own risk. If any code samples or other technology this work contains or describes is subject to open source licenses or the intellectual property rights of others, it is your responsibility to ensure that your use thereof complies with such licenses and/or rights. 979-8-341-60763-7 [FILL IN]
Page
5
Brief Table of Contents (Not Yet Final) Part 1: Foundations of SRE Chapter 1: Introduction (unavailable) Chapter 2: The Current Production Environment (unavailable) Part 2: The Sociotechnical Landscape of SRE Chapter 3: The Value of Reliability (unavailable) Chapter 4: The Cultural Context of SRE Chapter 5: Modeling What We Do and Why (unavailable) Chapter 6: Organizational Structures for SRE (unavailable) Chapter 7: SRE Engagement Models (unavailable) Part 3: Modern SRE Practices Chapter 8: SLOs (unavailable) Chapter 9: Observability and Monitoring (unavailable) Chapter 10: Incident Management and On-Call (unavailable) Chapter 11: Learning From Incidents Chapter 12: Automation and Tooling (unavailable) Chapter 13: Testing for Reliability in Production (unavailable) Chapter 14: Safety Engineering for Software: Beyond Reacting to Failure (unavailable) Chapter 15: Capacity (unavailable) Chapter 16: Software Engineering in SRE (unavailable)
Page
6
Chapter 17: SRE for AI (unavailable) Chapter 18: AI for SRE (unavailable) Chapter 19: Data Management (unavailable) Part 4: Management and Future of SRE Chapter 20: SRE Career Management and Training (unavailable) Chapter 21: SRE in Diverse Environments/Other Industries (unavailable) Chapter 22: The Future of SRE (unavailable)
Page
7
Chapter 1. The Cultural Context of SRE A NOTE FOR EARLY RELEASE READERS With Early Release ebooks, you get books in their earliest form—the author’s raw and unedited content as they write—so you can take advantage of these technologies long before the official release of these titles. This will be the 4th chapter of the final book. Please note that the GitHub repo will be made active later on. If you’d like to be actively involved in reviewing and commenting on this draft, please reach out to the editor at jleonard@oreilly.com. Organizational culture is foundational to the long-term success or failure of an SRE transformation. While SRE relies on technical practices and processes, these are deeply influenced by the people and environment surrounding them. We aren’t saying that having a conducive culture on Day 1 is a prerequisite, but rather that building a supportive culture is helpful in establishing and maintaining a successful SRE practice. In this chapter, we discuss data-driven, research-backed arguments for the critical importance of culture. This chapter will help you identify cultural elements that support SRE, articulate why culture is crucial for SRE success, build a plan to establish and foster a culture of reliability, and recognize and mitigate cultural antipatterns. The Case for Culture
Page
8
Let’s start by laying out the essential case for why culture is a powerful catalyst for achieving high reliability for our users, happy engineers, and overall organizational success in technology-driven environments. “The fifth nine is people.” —Trisha Weir, Google SRE Manager The "nines" in this quote refer to the number of nines in a reliability percentage, and 99.999% availability is referred to as "five nines“. This perspective emphasizes that while you can improve reliability through processes and tools, human behavior, collaboration, and the underlying culture are indispensable for reaching and sustaining elite performance. Once this status is reached, you’ll unlock improved reliability by up to an order of magnitude. Ben Treynor Sloss, the founder of SRE at Google, views the ability for SRE teams to be a dedicated voice for reliability as critically important, and cultural elements are essential for making SREs effective. Let’s explore this further. Key Elements of a Reliability-Focused Culture At its core, SRE applies a software engineering mindset to managing services in production. This approach fosters structured, strategic, and reproducible improvements, enabling systems to scale sustainably while reducing reliance on manual effort. This engineering mindset forms the bedrock upon which a robust reliability culture thrives. What practical steps are needed to cultivate a reliability-focused culture and turn these principles into day-to-day reality? A high-performing culture focused on reliability is built upon several key elements that prioritize alignment, learning, collaboration, and user satisfaction. These cultural components are crucial for enabling teams to effectively manage complex systems, respond to incidents, and drive continuous improvement. Establishing these elements provides the necessary environment for SRE practices to thrive and deliver meaningful results.
Page
9
Let’s start with the basics. There are two fundamental aspects of a reliability-focused culture: Focus on the user SRE is inherently user-centric. It involves a deep understanding of end users’ needs and goals, and leveraging their signals to continuously improve products and services. The SRE approach is not reliability at any cost. Instead, SRE focuses on understanding the end users’ and the business’s expectations on the reliability/velocity spectrum. Maintaining reliable systems is fundamentally about meeting these expectations, which in turn keeps users happy. Measuring system health and performance closely aligned to the user experience helps ensure their satisfaction. Reliability as a core value The organization committing to reliability is a prerequisite. Reliability is understood not just as a technical goal, but as an essential feature of any product or service. Reliability needs to be valued, recognized and rewarded across the organization, from engineers to executives, ensuring SRE has a seat at the decision-making table. Next, we’ll examine seven cultural elements that foster both reliability as a core value and focus on the user, as well as the high standards of excellence for which SRE is known. These are: A data-driven mindset Collaboration and shared responsibility Learning and continuous improvement Blamelessness
Page
10
Psychological safety Trust Empowerment Fostering a Reliability-Focused Culture In the previous section, we defined several dimensions that represent the elements of a reliability-focused culture. The two fundamental aspects, Reliability as a Core Value and Focus on the User, underpin everything else. Now let’s deep dive on how the remaining seven cultural aspects foster those two. We’ll look across the software development lifecycle at how organizations can integrate reliability-focused culture and practices into their business. Data-driven mindset SRE requires a data-driven approach to prioritize efforts and minimize future incidents. You need to establish different metrics–for example, to measure how well you’re doing against your targets, alert when there is a problem, and debug problems to find root causes. Adopting a measurement culture facilitates this data-driven approach. In order to set the right goals, allow for transparency, and build data-driven decision making processes, you need to measure everything important or relevant—not just reliability metrics, but also toil metrics and data that represents the customer experience. This data helps you identify gaps and prioritize solutions based on business needs and customer expectations. For example, the data from postmortems (also known as post-incident reviews or PIRs) is crucial for identifying patterns and guiding investments in failure analysis. SRE was created to handle the operational and data growth for Google. This scaling required not just technical changes, but also organizational and human changes. We had to align the organization to a shared goal–happy customers–without hindering innovation. To achieve this goal, we also had
Page
11
to align incentives across different organizations–business, development, and operations. This requires a common language, spanning the product lifecycle. In the field of system reliability, that common language is now widely established through the concepts of Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets. These elements of measurement form the bedrock of an SRE measurement culture and contextualize our data-driven mindset. We will discuss these fundamental SRE concepts in more detail in Part 3. To innovate rather than be fully occupied with the unachievable goal of “100% reliable,” we needed the ability to numerically estimate the risk ceiling and failure tolerance the system can have and still meet user expectations. We express this as an error budget–the amount of downtime our users are willing to tolerate. We measure how we are doing against our error budget, and use this for data-driven processes, decision making, and work planning. For example, we might decide to freeze rollouts, or to temporarily reassign part of a team to reliability work. Measurement culture applies not only to reliability metrics, but to every aspect of our business which underpins the high standards of excellence to which we hold ourselves. For example, we have a very strong postmortem culture, including detailed postmortem reports and tracking of postmortem action items. When we noticed systemic delays in writing the postmortem reports themselves and implementing postmortem action items, we took a data-driven approach. We collected metrics and set SLOs related to postmortem processes, such as the maximal time to write the first draft of a postmortem and to review and publish it. We also created dashboards, alerts, and management-visible reporting to track adherence to this SLO. This led to a reduction in report publication time, increased implementation of action items, and a decrease in recurrent “same root cause” incidents. Similarly, we introduced SLOs (with reporting and clear expectations) for handling high priority internal issues, particularly the time to provide internal status updates and the time to resolve the issue. For incidents, we set and enforce goals for metrics such as:
Page
12
Percentage of unactionable incidents. A high rate of unactionable incidents points to noisy or poorly configured monitoring. Percentage of manually detected incidents. A high rate of problems first reported by users or other teams suggests incorrectly defined SLIs or inadequate monitoring coverage. Page per incident ratio. Too many alerts for a single incident might distract an oncaller and make it harder to diagnose what’s going on. We apply the same measurement culture to people operations. To prevent burnout, we track metrics for on-call time. We also track how frequently on- callers have to apply manual changes in production, and we strive to keep this number as close to zero as possible. Ultimately, a data-driven mindset is the basis for everything SREs do, from measuring performance to driving effective prioritization and collaboration. Collaboration and shared responsibility The broad scope of SRE responsibility requires tight collaboration and alignment, which must occur at all levels: within the SRE team itself, within the larger SRE organization, and with the relevant software development teams. Within a team We often describe silos between functions or organizations, but silos can also exist within teams in the form of isolation of information, expertise, or skills. Common reasons for formation of knowledge silos are hoarding knowledge for job security, ignorance (not understanding why collaboration is beneficial), distrust (not believing that the potential partnership would actually be helpful), or very often just a lack of time or clear expectations. To overcome a tendency toward siloization, it is important to create a culture where people benefit from sharing information. Learning-focused organizations encourage a growth mindset, investing in continuous improvement rather than expecting people to know everything. We strive to
Page
13
actively reward collaboration through recognition and promotion. To achieve this, we encourage transparency and broad knowledge sharing through various strategies. One mechanism to share updates and review reliability issues is a weekly production meeting. In this meeting, SRE team members update each other on previous incidents and learnings, as well as upcoming production changes, launches and tests. Another strategy is weekly training sessions, where different team members present on products, technology, SRE fundamentals, personal initiatives, etc. In addition to knowledge sharing, training sessions enable personal growth through teaching, improving on-caller preparation, and encouraging a culture of learning. A third format for sharing knowledge is a “Wheel of Misfortune” (i.e., “Tabletop Exercise” or “Walk the Plank”), where team members can practice resolving different incidents and learn from experts. We strive to avoid relying on the heroics of individuals who, when needed, make extraordinary efforts on their own to solve or prevent complex issues. While often satisfying and subject to praise in the moment, this behavior is bad in the long run and leads to several SRE anti-patterns: It masks systemic problems, which never get fixed. The team isn’t motivated to set a realistic SLO, or to work on long-term systemic improvements. It cultivates a team culture that reinforces unrealistic expectations about the work expected from team members It’s unsustainable and leads to burnout. It leads to siloing. Within the broader SRE organization Junior SREs benefit from interacting with more experienced SREs across the organization to learn methodologies and best practices. They can also
Page
14
observe and replicate experienced SREs’ personal and organizational growth strategies. Where possible, geographical co-location positively contributes to this learning. Similar to team-level production meetings, Google has site-wide weekly ops reviews, where teams share on-call challenges and learnings. This includes previously handled incidents, new project designs, and new automation beneficial to SREs. Similarly, Ben Treynor Sloss regularly publishes a company-wide report titled “Google’s Greatest Hits & Misses”: a transparent and easily readable quarterly review of the most severe user or revenue impacting incidents, aiming to teach engineers how to prevent outages and build reliable systems. Tech Talks where subject matter experts in one team give a presentation or demo to the broader organization also foster collaboration, shared responsibility, and the exchange of knowledge. We’ll dig into this topic more in our chapter on hiring, training, and career management. Outside SRE Google intentionally separates SRE from the development function, enabling SRE autonomy–SRE teams don’t report to Dev leads. The goal is to have close collaboration between SRE and Dev as equal partners. A user-focused culture provides a common framework for SRE and Development to maintain appropriately reliable systems, where appropriate reliability is defined by customer needs. Business owners have to define the “right” level of reliability for a given product, and the organizational culture should enable both the developers and operators to enforce this level. The SRE and development organizations at Google are intentionally structured to balance two critical priorities. The primary focus of Development is to deliver features that contribute to the product’s functionality, and the primary focus of SRE is to deliver reliability (itself a mission critical feature). Both work toward the common goal of creating a valuable and stable product, and both roles must feel a sense of product ownership and align decisions and priorities to business goals. For example, a well established product requiring five 9s of reliability will have
Page
15
purposely slow rollouts with automatic rollbacks and extensive canarying. Conversely, with new products, customers are often more interested in seeing new features fast and may be more tolerant to outages, so the Dev and SRE teams might decide on a lower service level objective to shift the balance towards velocity and innovation. This does not mean that Development can/should unilaterally ignore reliability or brush SRE considerations aside due to the fast-moving nature of the emerging product. Instead, Dev and SRE should work together with clear agreement on the trade-offs being made and why. SLOs and error budgets allow business, dev, and ops teams to track progress toward agreed-upon business goals. We all want to balance development velocity with reliability in a way that delivers maximum value to our users. Collaboration between SREs and Devs is further driven by Dev participation in production meetings, collaboration on incident and postmortem reviews, and SRE-Dev leadership operational reviews, where the teams jointly review production metrics and leading indicators. A frequent point of contention is how early SRE should be involved in the design of a new product. Late involvement can lead to SRE reviews delaying launches if the risks raised about designs or operations don’t meet reliability standards. As reliability experts, SREs must be able to consult on the reliability and scalability of future products. Also, as production owners, SREs must be aware of new developments ahead of time, to allocate human and compute resources needed for productionization and support. Learning and continuous improvement Never Let a Good Crisis go to Waste. —widely attributed to Winston Churchill Innovation, such as adding new features, inherently involves change and risk, making some level of system failure inevitable. We see outages as valuable opportunities to learn and improve, but this is possible only if we identify the procedural and/or systematic causes. Postmortems are a
Page
16
formalized process for learning from outages. We internally publish a written record of the incident, including mitigation actions, impact, root cause(s), and follow-up actions to prevent recurrence. Postmortems provide a broad-based mechanism to prevent and reduce recurring incidents or at least reduce their likelihood and impact. Additionally, we share postmortems widely to help teams collectively learn from outages, translating lessons into concrete action for improvement, and enabling others to prevent similar incidents in their own systems. Good postmortems are blameless, preventing conversations about who’s at fault. Rather than seeking to blame, the focus is on understanding and improvement. The postmortem process is just one example of a mechanism to drive continuous improvement and underpin the high standards of excellence that SRE engenders. At its core, learning and continuous improvement is about leveraging SRE’s data-driven and software engineering mindset to fundamentally improve the systems themselves to make them more autonomous and to prevent future outages. This thread around continuous improvement is woven throughout a number of chapters including Learning from Incidents, STPA and CAST, Hiring, Training, and Career Management, and even Automation and Tooling. Blamelessness and blame-awareness Blameless postmortems are one aspect of a blameless culture. Creating a safe environment and fostering a culture of learning requires understanding blame and acknowledging cognitive biases. Blame can be a way to discharge discomfort, as Brené Brown suggests. It’s crucial to recognize cognitive biases such as the Fundamental Attribution Error (where we attribute others’ actions to internal traits and our own to external circumstances) and Hindsight Bias (the tendency to see past events as predictable). Blamelessness shifts the focus of responsibility from people to systems and processes after an incident. Finger-pointing leads to risk aversion, fear, and a tendency to hide mistakes, which reduces situational awareness, increases
Page
17
the risk of problems accumulating, and limits the organization’s ability to be proactive–stunting the ability to innovate and improve. Instead, a blameless culture focuses blame on the system and processes that allowed the incident to happen. Blameless culture assumes that individuals act in good faith and make decisions based on the best information available. Blame awareness takes this a step further. It acknowledges that blame is a natural human reaction, but instead of ignoring it, it encourages a conscious and curious examination of why an action that seems like a mistake made sense to the person who took it at the time. It asks, “Given the circumstances, the available information, and the pressures of the moment, why was this a reasonable thing to do?” This approach seeks to understand the “why” behind the “what,” leading to a deeper understanding of systemic flaws. This shift from blame to understanding, however, should not be treated as zero-accountability. While postmortems do not focus on human errors committed during an outage, it is important that everyone is held to high engineering standards. We have expectations for everyone in Google engineering, and these are especially critical when working on systems which have a wide scope of impact to Google’s users and its business. Postmortems should assess whether appropriate engineering standards were observed in both code and system operation, and to prescribe appropriate follow-ups where deficiencies are discovered.
Page
18
NOTE What if an employee acts maliciously? A common concern that we hear from those aspiring to adopt SRE principles, particularly in finance and regulated industries, is: what if an employee undertakes a malicious action (e.g., stealing)? Isn’t blame or holding people strictly accountable appropriate in that case? The principles of Just Culture can distinguish between honest mistakes, at-risk behavior, reckless conduct, or plain criminal cases. The latter two are rare cases where disciplinary action or legal consequences are appropriate. However, when referring to day-to-day employee actions we want to avoid blame to encourage appropriate risk taking and innovation. Even in cases that cross a line, it is important to do a postmortem and ask: what can we improve in our systems to make it impossible for someone to cross that line? Are we using the principle of least privilege? Do we have auditing enabled on those resources? Do sensitive operations have appropriate review and approval processes? Psychological safety According to Dr. Amy Edmondson, psychological safety is a belief that one will not be punished or humiliated for speaking up with ideas, questions, concerns, or mistakes. A blameless culture is crucial for establishing psychological safety. With psychological safety, people feel comfortable escalating incidents early, reporting problems, asking questions to identify root causes, and sharing information. Without it, incident response might be delayed, resolution might take longer, and both learning and innovation will suffer. Moreover, psychological safety is a critical enabler of a culture of continuous improvement. When team members feel safe, they are empowered to question the status quo and suggest better ways of working, far outside the context of outages and postmortems. This manifests in questions like: “Why do we have to do this toilsome task this way? Does the system architecture make sense? Is it okay that we haven’t tested our backups? What if…?” Managers and organizations play a key role in creating an environment where people feel safe and empowered to voice ideas. This includes
Page
19
psychological safety around acknowledging one’s own fallibility (e.g. by saying “I didn’t understand”), modeling curiosity by asking questions, leading by example, encouraging others, and not discouraging questions from others (e.g., through bad habits like feigned surprise). Managers are explicitly responsible for their team’s psychological safety by framing work as a learning opportunity and not just an execution problem. People must feel confident that asking questions or requesting help will be rewarded, and not punished. Trust Trust is a fragile thing — hard to earn, easy to lose. —M.J. Arlidge To freely speak up about mistakes, people need trust in their team, their manager, and the organization–that their input will be used for system improvement, and not for blame or punishment. The following paragraphs describe several strategies that can help leaders establish trust between people in their organization. One strategy is to create an organization where people choose to operate from the “assume good intentions, where possible” approach. This enables people to safely assume that their colleagues are competent and trustworthy. Note that a culture that assumes employees are competent and trustworthy doesn’t mean you never address performance issues. However, it does mean that you evaluate performance objectively, without initial assumptions. Multiple studies suggest you need an experience of at least 5-7 positive exchanges to build an immunity from the trust damage that a single tense incident can cause (Gottman & Levenson, 1992; Losada & Heaphy, 2004; Sabey et al., 2019). We try to have many positive interactions to build the “assume good intentions” mindset, so when a later disagreement or a conflict occurs, people still remember the good side of their colleague, don’t jump into negative assumptions, and resolve the disagreement in a healthy and constructive way. For globally-distributed SRE teams, targeted
Page
20
in-person meetings (in addition to virtual interactions) can be a crucial investment in this trust reserve. Leaders at various levels enable trust by giving trust to their reports, acknowledging their own mistakes, and following up with accountability. Leaders should also model blamelessness, rewarding people for acknowledging and resolving issues. At the team level, managers can build trust and community through collaboration activities and exercises. Effective, work-focused approaches can include workshops to define team- specific communication styles and norms, or collaborative exercises like practicing incident response scenarios together. These activities help team members understand how their colleagues approach problems and communicate under pressure, building rapport and mutual respect through shared professional experiences. These approaches also improve psychological safety, performance and retention. Empowerment It’s important to empower SREs to speak up when reliability is threatened by a poor design or launch plan, even before metrics like error budget are impacted. SREs also need to regulate their operational workload. When system performance is not meeting expectations, they need to advocate for spending more time on reliability, even if it means slowing down feature releases. To prevent burnout from SRE teams getting buried in operational tasks, it’s vital to set dedicated time for reliability work during the development cycle, and not treat this work as an afterthought. This work helps SREs advance their careers and deliver value to the business. Empowerment goes even further than this. SREs must be in a position to understand and fix reliability issues in a system. SRE is also a culture of initiative. SREs are not only allowed, but expected to drive continuous improvement. At Google, one way this empowerment is manifested is through access to the code and the ability to fix things in the codebase. Any member of the